Skip to main content

OpenAI-compatible endpoints

One provider for every endpoint that speaks POST {api_base}/chat/completions and does not have a dedicated plugin of its own: Groq, DeepSeek, Mistral, Together, Fireworks, Perplexity, vLLM, LM Studio, LiteLLM proxies, Azure OpenAI.

Two decisions make it generic rather than merely configurable:

  • The registry id is renamable. Point it at Groq and set id = "groq", and /async status, the dashboard and every metric label say Groq instead of a placeholder.
  • Nothing about the endpoint is assumed. api_base and model are required — there is no honest default across ten vendors. The API key is optional (vLLM and LM Studio take none). JSON mode is switchable off (some endpoints reject response_format). Extra headers pass through verbatim.
danger
The section key is always compat

The config section must be [llm.providers.compat], and [llm] provider must say "compat". The id key inside the section renames what the provider is called at runtime — it does not rename the section.

[llm]
provider = "compat" # ← not "groq"

[llm.providers.compat] # ← not [llm.providers.groq]
id = "groq" # ← this is what shows up in /async status

An unknown id fails at startup with a message listing what the binary does support. A consequence worth knowing: one compat provider per deployment — two of them cannot both occupy the single compat section key.

Configuration keys​

KeyRequiredDefaultNotes
api_base✅—Absolute http(s) endpoint root. A query string is preserved (Azure needs it)
model✅—No cross-vendor default exists
id—openai-compatRegistry id; lowercase [a-z0-9-], 1–32 chars
display_name—derived from idHuman label in status output
api_key_env——Omit for local servers that take no credential
max_tokens—512Must be a positive integer; a bad value is a config error, not a silent fallback
json_mode—trueSet false for endpoints that reject response_format
[….headers]—emptyTable of headers sent verbatim on every request

Auth order: the bearer token from api_key_env is applied first, then your headers — so a headers entry can replace Authorization for a gateway that expects its own scheme.

warning
headers values are read from the file verbatim

Anything you put in [llm.providers.compat.headers] is a literal string in the config file — it is not an *_env indirection. Prefer api_key_env wherever the endpoint accepts Authorization: Bearer; that path keeps the credential in the environment where it belongs. Use headers for a secret only when the endpoint gives you no alternative, and treat the config file as sensitive when you do.

Groq​

Fast inference, generous free tier — a good cheap classifier.

[llm]
provider = "compat"

[llm.providers.compat]
id = "groq"
api_base = "https://api.groq.com/openai/v1"
api_key_env = "GROQ_API_KEY"
model = "llama-3.3-70b-versatile"

DeepSeek​

[llm]
provider = "compat"

[llm.providers.compat]
id = "deepseek"
display_name = "DeepSeek"
api_base = "https://api.deepseek.com/v1"
api_key_env = "DEEPSEEK_API_KEY"
model = "deepseek-chat"
max_tokens = 512

Mistral​

[llm.providers.compat]
id = "mistral"
api_base = "https://api.mistral.ai/v1"
api_key_env = "MISTRAL_API_KEY"
model = "mistral-small-latest"

vLLM (self-hosted, no credentials)​

[llm.providers.compat]
id = "vllm"
display_name = "vLLM (local)"
api_base = "http://192.168.1.50:8000/v1"
model = "Qwen/Qwen3-8B"
# Older builds reject `response_format`; the classifier parser tolerates
# fenced and prose-wrapped JSON, so turning it off stays safe.
json_mode = false

Note there is no api_key_env — a bare vLLM server authenticates nobody, and the provider sends no Authorization header when no key is configured.

LM Studio​

LM Studio's local server is OpenAI-compatible on :1234:

[llm.providers.compat]
id = "lmstudio"
display_name = "LM Studio"
api_base = "http://localhost:1234/v1"
model = "qwen2.5-7b-instruct"
json_mode = false

Azure OpenAI​

Azure needs three things the others do not: a deployment URL, a pinned api-version query parameter, and the api-key header instead of a bearer token. All three work here — a query string on api_base is preserved onto every request.

[llm.providers.compat]
id = "azure-openai"
api_base = "https://ACCOUNT.openai.azure.com/openai/deployments/DEPLOYMENT?api-version=2024-10-21"
model = "gpt-4o-mini"

[llm.providers.compat.headers]
"api-key" = "..."

If your Azure endpoint accepts Authorization: Bearer instead, prefer the dedicated OpenAI provider — it keeps the credential in the environment.

json_mode: when to turn it off​

The classifier asks for response_format: {"type":"json_object"}. A rewrite never does — it is prose the author posts verbatim.

Turn json_mode off when the endpoint rejects the field outright (older vLLM builds, some proxies) and every classify call is failing with a 400. It is safe to do so: the classifier parser already tolerates fenced and prose-wrapped JSON. What it will not do is invent a verdict from a reply that does not contain one — a malformed reply is an error, and an error is read as "no opinion".

Same prompts, same verdicts​

Prompts and reply parsing come from the shared crate, so this provider judges messages exactly like the dedicated ones do. A chain of compat(groq) → openai → ollama changes who answers, never what was asked.

Diagnosing a new endpoint​

SymptomLikely cause
Startup: ``api_base is required and must be an absolute http(s) endpoint rootmissing scheme, or you gave a path without http://
Startup: ``model is requiredno cross-vendor default exists — name the model
Startup: unknown llm provideryou used [llm.providers.groq]; the section key must be compat
Every classify fails with 400the endpoint rejects response_format → json_mode = false
401 on a local serveryou set api_key_env for a server that takes no credential
Replies truncate mid-JSONmax_tokens too low; 512 is the default for a reason

Errors carry the status and a body excerpt — never a credential. everasync doctor will not exercise the LLM chain (it health-checks channels and connectors), so the fastest test is to set sensitivity = "balanced", post an ambiguous message, and watch ever_async_llm_calls_total on /metrics.