Skip to main content

Gemini

Google Gemini over the Generative Language API. Gemini is not OpenAI-shaped, so none of the shared chat-completions helpers apply — but every prompt and every parse still come from the shared crate, because one wording across providers is the whole point of the fallback chain.

Good for​

  • Teams already on Google Cloud, or already holding a Google AI Studio key.
  • A very cheap, very fast classifier: the Flash family is exactly the shape this job wants — high volume, low difficulty.
  • Long context, if your channel messages are unusually large (they rarely are).

Configuration​

[llm]
provider = "gemini"

[llm.providers.gemini]
api_key_env = "GEMINI_API_KEY" # required
model = "gemini-2.5-flash" # with or without the "models/" prefix
max_tokens = 512
# api_base = "https://generativelanguage.googleapis.com/v1beta"
export GEMINI_API_KEY=...
KeyRequiredDefaultNotes
api_key_env✅—Env var holding the key
model—gemini-2.5-flashA models/ prefix is accepted and stripped
api_base—https://generativelanguage.googleapis.com/v1betaTrailing slash trimmed
max_tokens—512A negative or oversized value falls back to the default rather than failing startup

The model name is part of the URL path on this API (POST /models/{model}:generateContent), not of the request body — which is why a catalog id pasted back verbatim (models/gemini-2.5-flash) is accepted and normalized instead of producing a double prefix.

The API key travels in a header​

x-goog-api-key: <key>

Never as a ?key= query parameter. A secret in a query string leaks into proxy logs, request traces, browser history and every error report the moment anything goes wrong. If you put a proxy in front of this provider, make sure it does not rewrite the credential into the URL.

Choosing a model​

WantTry
Cheapest, fastest — the right defaulta Flash model (built-in default: gemini-2.5-flash)
Better judgement on borderline messagesa Pro model

list_models reads Google's catalog and keeps only models that advertise the generateContent method — embedding and legacy models advertise other methods and cannot answer a prompt, so offering them in the dashboard picker would only produce broken configs. Catalog ids come back namespaced (models/…); the prefix is stripped before storage because the request path supplies it.

Both calls run at temperature 0.2, fixed. Both jobs want the model's most predictable answer, not its most creative one: a classifier must be reproducible, and a rewrite must stay the author's message rather than become the model's.

Cost and privacy​

  • Only ambiguous messages are sent, and only under balanced / high. Under conservative nothing is sent at all.
  • What is sent: the message text and the team charter; for a rewrite, resolved context items too.
  • Free-tier Google AI Studio keys may be used to improve Google's products — check the terms that apply to your key before pointing this at real work conversations. A paid Google Cloud project has different terms.
  • Requests time out after 30 seconds and fall through to the next provider.
  • If message content must not reach a vendor at all, use Ollama.

Errors and secrets​

Blocked and empty replies are errors, never guesses. Gemini can return a well-formed response with no usable candidate — a safety block, a stop reason that truncated the answer, an empty parts array. Each of those is an AsyncError::Llm, which the pipeline reads as "no opinion"; the heuristic verdict stands and nothing is invented.

Error messages carry the status and a capped body excerpt — never the API key, which exists only in the request header.

A 400 with an API_KEY_INVALID reason is the key; a 404 is usually a model id that is not available in your region or on your API version; a 429 is what the fallback chain exists for.