Gemini
Google Gemini over the Generative Language API. Gemini is not OpenAI-shaped, so none of the shared chat-completions helpers apply — but every prompt and every parse still come from the shared crate, because one wording across providers is the whole point of the fallback chain.
Good for
- Teams already on Google Cloud, or already holding a Google AI Studio key.
- A very cheap, very fast classifier: the Flash family is exactly the shape this job wants — high volume, low difficulty.
- Long context, if your channel messages are unusually large (they rarely are).
Configuration
[llm]
provider = "gemini"
[llm.providers.gemini]
api_key_env = "GEMINI_API_KEY" # required
model = "gemini-2.5-flash" # with or without the "models/" prefix
max_tokens = 512
# api_base = "https://generativelanguage.googleapis.com/v1beta"
export GEMINI_API_KEY=...
| Key | Required | Default | Notes |
|---|---|---|---|
api_key_env | ✅ | — | Env var holding the key |
model | — | gemini-2.5-flash | A models/ prefix is accepted and stripped |
api_base | — | https://generativelanguage.googleapis.com/v1beta | Trailing slash trimmed |
max_tokens | — | 512 | A negative or oversized value falls back to the default rather than failing startup |
The model name is part of the URL path on this API
(POST /models/{model}:generateContent), not of the request body — which is
why a catalog id pasted back verbatim (models/gemini-2.5-flash) is accepted
and normalized instead of producing a double prefix.
The API key travels in a header
x-goog-api-key: <key>
Never as a ?key= query parameter. A secret in a query string leaks into
proxy logs, request traces, browser history and every error report the moment
anything goes wrong. If you put a proxy in front of this provider, make sure it
does not rewrite the credential into the URL.
Choosing a model
| Want | Try |
|---|---|
| Cheapest, fastest — the right default | a Flash model (built-in default: gemini-2.5-flash) |
| Better judgement on borderline messages | a Pro model |
list_models reads Google's catalog and keeps only models that advertise the
generateContent method — embedding and legacy models advertise other methods
and cannot answer a prompt, so offering them in the dashboard picker would only
produce broken configs. Catalog ids come back namespaced (models/…); the
prefix is stripped before storage because the request path supplies it.
Both calls run at temperature 0.2, fixed. Both jobs want the model's most predictable answer, not its most creative one: a classifier must be reproducible, and a rewrite must stay the author's message rather than become the model's.
Cost and privacy
- Only ambiguous messages are sent, and only under
balanced/high. Underconservativenothing is sent at all. - What is sent: the message text and the team charter; for a rewrite, resolved context items too.
- Free-tier Google AI Studio keys may be used to improve Google's products — check the terms that apply to your key before pointing this at real work conversations. A paid Google Cloud project has different terms.
- Requests time out after 30 seconds and fall through to the next provider.
- If message content must not reach a vendor at all, use Ollama.
Errors and secrets
Blocked and empty replies are errors, never guesses. Gemini can return a
well-formed response with no usable candidate — a safety block, a stop reason
that truncated the answer, an empty parts array. Each of those is an
AsyncError::Llm, which the pipeline reads as "no opinion"; the heuristic
verdict stands and nothing is invented.
Error messages carry the status and a capped body excerpt — never the API key, which exists only in the request header.
A 400 with an API_KEY_INVALID reason is the key; a 404 is usually a model id
that is not available in your region or on your API version; a 429 is what the
fallback chain exists for.