Skip to main content

AI providers

The AI layer is optional. Turn it off entirely and Ever Async still catches naked greetings, contentless help pleas and artifact-without-link messages — deterministically, in microseconds, for zero tokens.

What a provider adds is judgement on the ambiguous middle: the short message with no links and no strong pattern that a regex should not be allowed to guess about.

What a provider is asked to do​

Exactly two things, and never anything else:

OperationWhenOutput
classifythe heuristics were not certain and sensitivity is balanced or highstrict JSON: verdict, confidence, what is missing, reasons
suggest_rewritea nudge is about to be sentprose — the message the author could have written

Both prompts, the strict-JSON parsing of the classifier reply and the rewrite cleanup live in one shared crate (ever-async-llm-shared), not in the provider crates. That is deliberate:

A provider that invented its own prompt would make the fallback chain change its mind about a message purely by moving down a rung — the one thing the chain must not do.

A provider crate is therefore only its HTTP shape: an endpoint, an auth header, and a response mapping.

Malformed replies are errors, never guesses​

If a model answers something that is not the demanded shape, the parser fails. The pipeline reads an LLM error as "no opinion" and falls back to heuristics. A lenient parser that salvaged half a bad reply would turn a model outage into a wave of false positives — see sensitivity.

Configuration shape​

One primary, an ordered list of fallbacks, and one section per provider:

[llm]
provider = "openrouter" # primary — omit the whole [llm] block to run heuristics-only
fallback = ["openai", "ollama"] # tried in order when the one before errors

[llm.providers.openrouter]
api_key_env = "OPENROUTER_API_KEY"
model = "anthropic/claude-sonnet-5"

[llm.providers.openai]
api_key_env = "OPENAI_API_KEY"
model = "gpt-5-mini"

[llm.providers.ollama]
model = "llama3.1:8b"

Rules that are easy to get wrong:

  • The section key is the provider id. [llm.providers.openrouter] configures the provider named openrouter. Naming a provider in [llm] without giving it a section is a hard startup error — you asked for something the binary cannot deliver.
  • fallback is deduplicated against the primary, so listing the primary again is harmless.
  • An unknown id fails at startup, listing what the binary does support. That covers both a typo and a provider whose cargo feature was compiled out.
  • Omit [llm] entirely and the LLM layer is simply off.

The fallback chain​

The chain is presented to the pipeline as a single provider. The pipeline does not know it exists.

Each call is tried against every provider in order and the first success wins. Every failure is logged and counted before the next is tried:

ever_async_llm_calls_total{provider="openrouter",operation="classify",outcome="error"} 3
ever_async_llm_calls_total{provider="openai",operation="classify",outcome="ok"} 3

If every provider fails, the last error is returned — and the pipeline treats that as "no opinion". A dead AI layer degrades the product, it never breaks it.

This is what a fallback is for: rate limits, an outage, a spend cap. A local Ollama makes an excellent last rung — slower, but it answers with the internet unplugged.

note
list_models reads the primary only

The dashboard's model picker calls list_models, and the router forwards that to the first provider in the chain. Fallbacks are for answering, not for browsing catalogs.

Choosing a provider​

ProviderReach for it whenPage
OpenRouteryou want to change models without changing keys, and want a live catalog + provider routingOpenRouter
OpenAIyou already have an OpenAI account, or you are on Azure OpenAIOpenAI
Anthropicyou want Claude directly, on the Messages APIAnthropic
Geminiyou are on Google Cloud / already hold a Google AI keyGemini
Grokyou have an xAI accountGrok
Ollamamessage text must not leave your network, or you need an offline last resortOllama
OpenAI-compatibleGroq, DeepSeek, Mistral, Together, Fireworks, vLLM, LM Studio, LiteLLM, AzureOpenAI-compatible

Picking a model​

Two jobs, different shapes:

  • Classification is high-volume and low-difficulty. It runs on every ambiguous message. Every provider defaults to a small fast model for exactly this reason.
  • A rewrite is one short message, not an essay. The token cap defaults to 512 across every provider.

So the cost lever is the classification model, and the quality lever is whichever model you would trust to judge whether a colleague's message is answerable. If a wrong nudge costs your team more than a token does, pay for the better model — that is why the OpenRouter default is Sonnet rather than something cheap.

Set model explicitly. Inheriting a default means the model changes under you when the default does.

Cost and privacy at a glance​

Cost per messageMessage text leaves your network
No provider (conservative)zeronever
balanced / high + hosted provideronly for ambiguous messagesyes — to that vendor
balanced / high + Ollamazero (your hardware)never

Only the ambiguous middle is ever sent — a clear pass and a textbook fail are both decided locally. Under conservative, nothing is ever sent at all.

What is sent​

For classify: the message text and the team charter. For suggest_rewrite: the same, plus any context items that resolved.

The charter is what makes the AI layer team-specific — the model judges against your rules, not a generic idea of good writing. Change the charter and the verdicts change with it.

Verifying what is running​

/async status
Ever Async here is *ON* — sensitivity `Balanced`, nudge delay 150s, …
AI: OpenRouter (+2 fallback) · `anthropic/claude-sonnet-5`

or over HTTP:

curl -s localhost:8100/api/v1/status | jq '.plugins[] | select(.kind=="llm")'

everasync doctor reports each member of the chain individually — the router presents itself as one provider, so doctor is the way to see all the rungs.