AI providers
The AI layer is optional. Turn it off entirely and Ever Async still catches naked greetings, contentless help pleas and artifact-without-link messages — deterministically, in microseconds, for zero tokens.
What a provider adds is judgement on the ambiguous middle: the short message with no links and no strong pattern that a regex should not be allowed to guess about.
What a provider is asked to do
Exactly two things, and never anything else:
| Operation | When | Output |
|---|---|---|
classify | the heuristics were not certain and sensitivity is balanced or high | strict JSON: verdict, confidence, what is missing, reasons |
suggest_rewrite | a nudge is about to be sent | prose — the message the author could have written |
Both prompts, the strict-JSON parsing of the classifier reply and the rewrite
cleanup live in one shared crate (ever-async-llm-shared), not in the
provider crates. That is deliberate:
A provider that invented its own prompt would make the fallback chain change its mind about a message purely by moving down a rung — the one thing the chain must not do.
A provider crate is therefore only its HTTP shape: an endpoint, an auth header, and a response mapping.
Malformed replies are errors, never guesses
If a model answers something that is not the demanded shape, the parser fails. The pipeline reads an LLM error as "no opinion" and falls back to heuristics. A lenient parser that salvaged half a bad reply would turn a model outage into a wave of false positives — see sensitivity.
Configuration shape
One primary, an ordered list of fallbacks, and one section per provider:
[llm]
provider = "openrouter" # primary — omit the whole [llm] block to run heuristics-only
fallback = ["openai", "ollama"] # tried in order when the one before errors
[llm.providers.openrouter]
api_key_env = "OPENROUTER_API_KEY"
model = "anthropic/claude-sonnet-5"
[llm.providers.openai]
api_key_env = "OPENAI_API_KEY"
model = "gpt-5-mini"
[llm.providers.ollama]
model = "llama3.1:8b"
Rules that are easy to get wrong:
- The section key is the provider id.
[llm.providers.openrouter]configures the provider namedopenrouter. Naming a provider in[llm]without giving it a section is a hard startup error — you asked for something the binary cannot deliver. fallbackis deduplicated against the primary, so listing the primary again is harmless.- An unknown id fails at startup, listing what the binary does support. That covers both a typo and a provider whose cargo feature was compiled out.
- Omit
[llm]entirely and the LLM layer is simply off.
The fallback chain
The chain is presented to the pipeline as a single provider. The pipeline does not know it exists.
Each call is tried against every provider in order and the first success wins. Every failure is logged and counted before the next is tried:
ever_async_llm_calls_total{provider="openrouter",operation="classify",outcome="error"} 3
ever_async_llm_calls_total{provider="openai",operation="classify",outcome="ok"} 3
If every provider fails, the last error is returned — and the pipeline treats that as "no opinion". A dead AI layer degrades the product, it never breaks it.
This is what a fallback is for: rate limits, an outage, a spend cap. A local Ollama makes an excellent last rung — slower, but it answers with the internet unplugged.
list_models reads the primary onlyThe dashboard's model picker calls list_models, and the router forwards that
to the first provider in the chain. Fallbacks are for answering, not for
browsing catalogs.
Choosing a provider
| Provider | Reach for it when | Page |
|---|---|---|
| OpenRouter | you want to change models without changing keys, and want a live catalog + provider routing | OpenRouter |
| OpenAI | you already have an OpenAI account, or you are on Azure OpenAI | OpenAI |
| Anthropic | you want Claude directly, on the Messages API | Anthropic |
| Gemini | you are on Google Cloud / already hold a Google AI key | Gemini |
| Grok | you have an xAI account | Grok |
| Ollama | message text must not leave your network, or you need an offline last resort | Ollama |
| OpenAI-compatible | Groq, DeepSeek, Mistral, Together, Fireworks, vLLM, LM Studio, LiteLLM, Azure | OpenAI-compatible |
Picking a model
Two jobs, different shapes:
- Classification is high-volume and low-difficulty. It runs on every ambiguous message. Every provider defaults to a small fast model for exactly this reason.
- A rewrite is one short message, not an essay. The token cap defaults to 512 across every provider.
So the cost lever is the classification model, and the quality lever is whichever model you would trust to judge whether a colleague's message is answerable. If a wrong nudge costs your team more than a token does, pay for the better model — that is why the OpenRouter default is Sonnet rather than something cheap.
Set model explicitly. Inheriting a default means the model changes under you
when the default does.
Cost and privacy at a glance
| Cost per message | Message text leaves your network | |
|---|---|---|
No provider (conservative) | zero | never |
balanced / high + hosted provider | only for ambiguous messages | yes — to that vendor |
balanced / high + Ollama | zero (your hardware) | never |
Only the ambiguous middle is ever sent — a clear pass and a textbook fail are
both decided locally. Under conservative, nothing is ever sent at all.
What is sent
For classify: the message text and the team charter.
For suggest_rewrite: the same, plus any context items that resolved.
The charter is what makes the AI layer team-specific — the model judges against your rules, not a generic idea of good writing. Change the charter and the verdicts change with it.
Verifying what is running
/async status
Ever Async here is *ON* — sensitivity `Balanced`, nudge delay 150s, …
AI: OpenRouter (+2 fallback) · `anthropic/claude-sonnet-5`
or over HTTP:
curl -s localhost:8100/api/v1/status | jq '.plugins[] | select(.kind=="llm")'
everasync doctor reports each member of the chain individually — the router
presents itself as one provider, so doctor is the way to see all the rungs.