OpenAI-compatible endpoints
One provider for every endpoint that speaks POST {api_base}/chat/completions
and does not have a dedicated plugin of its own: Groq, DeepSeek, Mistral,
Together, Fireworks, Perplexity, vLLM, LM Studio, LiteLLM proxies, Azure OpenAI.
Two decisions make it generic rather than merely configurable:
- The registry id is renamable. Point it at Groq and set
id = "groq", and/async status, the dashboard and every metric label say Groq instead of a placeholder. - Nothing about the endpoint is assumed.
api_baseandmodelare required — there is no honest default across ten vendors. The API key is optional (vLLM and LM Studio take none). JSON mode is switchable off (some endpoints rejectresponse_format). Extra headers pass through verbatim.
compatThe config section must be [llm.providers.compat], and [llm] provider
must say "compat". The id key inside the section renames what the provider
is called at runtime — it does not rename the section.
[llm]
provider = "compat" # ← not "groq"
[llm.providers.compat] # ← not [llm.providers.groq]
id = "groq" # ← this is what shows up in /async status
An unknown id fails at startup with a message listing what the binary does
support. A consequence worth knowing: one compat provider per deployment —
two of them cannot both occupy the single compat section key.
Configuration keys
| Key | Required | Default | Notes |
|---|---|---|---|
api_base | ✅ | — | Absolute http(s) endpoint root. A query string is preserved (Azure needs it) |
model | ✅ | — | No cross-vendor default exists |
id | — | openai-compat | Registry id; lowercase [a-z0-9-], 1–32 chars |
display_name | — | derived from id | Human label in status output |
api_key_env | — | — | Omit for local servers that take no credential |
max_tokens | — | 512 | Must be a positive integer; a bad value is a config error, not a silent fallback |
json_mode | — | true | Set false for endpoints that reject response_format |
[….headers] | — | empty | Table of headers sent verbatim on every request |
Auth order: the bearer token from api_key_env is applied first, then your
headers — so a headers entry can replace Authorization for a gateway that
expects its own scheme.
headers values are read from the file verbatimAnything you put in [llm.providers.compat.headers] is a literal string in the
config file — it is not an *_env indirection. Prefer api_key_env
wherever the endpoint accepts Authorization: Bearer; that path keeps the
credential in the environment where it belongs. Use headers for a secret only
when the endpoint gives you no alternative, and treat the config file as
sensitive when you do.
Groq
Fast inference, generous free tier — a good cheap classifier.
[llm]
provider = "compat"
[llm.providers.compat]
id = "groq"
api_base = "https://api.groq.com/openai/v1"
api_key_env = "GROQ_API_KEY"
model = "llama-3.3-70b-versatile"
DeepSeek
[llm]
provider = "compat"
[llm.providers.compat]
id = "deepseek"
display_name = "DeepSeek"
api_base = "https://api.deepseek.com/v1"
api_key_env = "DEEPSEEK_API_KEY"
model = "deepseek-chat"
max_tokens = 512
Mistral
[llm.providers.compat]
id = "mistral"
api_base = "https://api.mistral.ai/v1"
api_key_env = "MISTRAL_API_KEY"
model = "mistral-small-latest"
vLLM (self-hosted, no credentials)
[llm.providers.compat]
id = "vllm"
display_name = "vLLM (local)"
api_base = "http://192.168.1.50:8000/v1"
model = "Qwen/Qwen3-8B"
# Older builds reject `response_format`; the classifier parser tolerates
# fenced and prose-wrapped JSON, so turning it off stays safe.
json_mode = false
Note there is no api_key_env — a bare vLLM server authenticates nobody, and
the provider sends no Authorization header when no key is configured.
LM Studio
LM Studio's local server is OpenAI-compatible on :1234:
[llm.providers.compat]
id = "lmstudio"
display_name = "LM Studio"
api_base = "http://localhost:1234/v1"
model = "qwen2.5-7b-instruct"
json_mode = false
Azure OpenAI
Azure needs three things the others do not: a deployment URL, a pinned
api-version query parameter, and the api-key header instead of a bearer
token. All three work here — a query string on api_base is preserved onto
every request.
[llm.providers.compat]
id = "azure-openai"
api_base = "https://ACCOUNT.openai.azure.com/openai/deployments/DEPLOYMENT?api-version=2024-10-21"
model = "gpt-4o-mini"
[llm.providers.compat.headers]
"api-key" = "..."
If your Azure endpoint accepts Authorization: Bearer instead, prefer the
dedicated OpenAI provider — it keeps the credential
in the environment.
json_mode: when to turn it off
The classifier asks for response_format: {"type":"json_object"}. A rewrite
never does — it is prose the author posts verbatim.
Turn json_mode off when the endpoint rejects the field outright (older
vLLM builds, some proxies) and every classify call is failing with a 400. It is
safe to do so: the classifier parser already tolerates fenced and
prose-wrapped JSON. What it will not do is invent a verdict from a reply that
does not contain one — a malformed reply is an error, and an error is read as
"no opinion".
Same prompts, same verdicts
Prompts and reply parsing come from the shared crate, so this provider judges
messages exactly like the dedicated ones do. A chain of
compat(groq) → openai → ollama changes who answers, never what was asked.
Diagnosing a new endpoint
| Symptom | Likely cause |
|---|---|
Startup: ``api_base is required and must be an absolute http(s) endpoint root | missing scheme, or you gave a path without http:// |
Startup: ``model is required | no cross-vendor default exists — name the model |
Startup: unknown llm provider | you used [llm.providers.groq]; the section key must be compat |
| Every classify fails with 400 | the endpoint rejects response_format → json_mode = false |
| 401 on a local server | you set api_key_env for a server that takes no credential |
| Replies truncate mid-JSON | max_tokens too low; 512 is the default for a reason |
Errors carry the status and a body excerpt — never a credential.
everasync doctor will not exercise the LLM chain (it health-checks channels
and connectors), so the fastest test is to set
sensitivity = "balanced", post an ambiguous message, and watch
ever_async_llm_calls_total on /metrics.