Skip to main content

OpenRouter

The provider to reach for when you want to choose a model. One key reaches every vendor, so the model id becomes the only thing you change when you want cheaper, faster or smarter — no new account, no new secret, no config migration.

Good for​

  • Trying Claude, GPT, Gemini and Llama against your own channel traffic without four accounts.
  • Teams that want the dashboard's model picker to be genuinely useful — it is fed live from OpenRouter's catalog.
  • Pinning which upstream provider serves a model, for latency, price or data-handling reasons.

Configuration​

[llm]
provider = "openrouter"
fallback = ["ollama"]

[llm.providers.openrouter]
api_key_env = "OPENROUTER_API_KEY" # required
model = "anthropic/claude-sonnet-5" # namespaced "vendor/model"
max_tokens = 512

# Attribution — OpenRouter shows these on its public app leaderboards
site_url = "https://docs-async.ever.co"
app_name = "Ever Async"

# Upstream routing (optional)
provider_order = ["anthropic", "google-vertex"]
allow_fallbacks = true
export OPENROUTER_API_KEY=sk-or-...
KeyRequiredDefaultNotes
api_key_env✅—Env var holding the key
model—anthropic/claude-sonnet-5Namespaced vendor/model
api_base—https://openrouter.ai/api/v1Trailing slash trimmed
max_tokens—512Clamped to 64–4096
site_url——Sent as HTTP-Referer
app_name—Ever AsyncSent as X-Title
provider_order—emptyArray of upstream provider slugs, tried in order
allow_fallbacks—unsetfalse pins routing to provider_order only

Model selection​

This is the reason to use OpenRouter, so it is worth doing deliberately.

Ids are namespaced: vendor/model.

model = "anthropic/claude-sonnet-5"
model = "openai/gpt-5-mini"
model = "google/gemini-2.5-flash"
model = "meta-llama/llama-3.3-70b-instruct"

The default is Sonnet, not something cheap, and that is a considered choice:

The ambiguous middle is a judgement call, and a wrong nudge costs more than a token.

Everything cheap and easy has already been decided by the heuristics before a model is consulted. What reaches OpenRouter is precisely the set of messages where being wrong is expensive — a false positive costs trust, and trust is the thing this product cannot buy back. See sensitivity.

If you want to cut cost anyway, the honest move is to leave the model alone and put the sensitivity back to conservative, which sends nothing at all.

The live catalog​

list_models reads GET /models and feeds the dashboard's picker directly. It is written to survive whatever the catalog omits — a missing name, price or context window drops that field, never the model. Prices are published per token and shown per million tokens, which is the unit people compare in.

Because the router forwards list_models to the primary provider only, put OpenRouter first if you want the catalog in the picker.

Provider routing​

provider_order pins which upstream actually serves the model:

# Prefer Anthropic's own capacity, then Vertex, and nothing else
provider_order = ["anthropic", "google-vertex"]
allow_fallbacks = false
  • provider_order — an array of upstream slugs, tried in order. Non-string and blank entries are dropped rather than forwarded as a routing rule nobody meant.
  • allow_fallbacks — leave it unset to keep OpenRouter's default (fallbacks allowed). Set false to hard-pin: the call fails rather than silently landing on an upstream you did not approve.

Reach for this when the choice of upstream matters — a data-handling agreement with one vendor, a latency difference between regions, or an inference provider you have benchmarked and the others you have not.

caution
allow_fallbacks = false narrows your safety net

Pinning removes OpenRouter's own redundancy. Pair it with an Ever Async fallback chain so a pinned upstream going down degrades to another provider instead of to no AI at all.

Attribution headers​

site_url (HTTP-Referer) and app_name (X-Title) are how OpenRouter attributes traffic on its public app leaderboards. Neither is a secret and neither affects routing — set them or do not.

Cost and privacy​

  • Only ambiguous messages are sent, and only under balanced / high.
  • Message text reaches OpenRouter and the upstream provider that serves the request. If that second hop is unacceptable, pin provider_order, or use Ollama and keep everything on your own hardware.
  • Replies are capped at max_tokens (512 by default); requests time out after 30 seconds.

Errors and secrets​

Errors carry the status and a capped body excerpt, never the key. A 401 is the key; a 404 on a valid key is almost always a mistyped namespace (claude-sonnet-5 instead of anthropic/claude-sonnet-5); a 402 means the account is out of credit — exactly the case the fallback chain is for.