OpenAI
Classification and rewrite suggestions over POST {api_base}/chat/completions.
The same crate drives Azure OpenAI and any OpenAI-compatible gateway — the
only difference is api_base.
Good for
- Teams that already have an OpenAI account and billing in place.
- Azure OpenAI deployments, where the model is a deployment name behind your own endpoint and the data-residency story is already signed off.
- A cheap, fast classifier:
gpt-5-mini-class models are the right shape for high-volume, low-difficulty judgement.
Configuration
[llm]
provider = "openai"
[llm.providers.openai]
api_key_env = "OPENAI_API_KEY" # required — names the env var; the key is never in the file
model = "gpt-5-mini" # set this explicitly
max_tokens = 512
# api_base = "https://api.openai.com/v1"
# organization = "org-..."
export OPENAI_API_KEY=sk-...
| Key | Required | Default | Notes |
|---|---|---|---|
api_key_env | ✅ | — | Env var holding the API key |
model | — | gpt-5-mini | Any chat model the account can reach |
api_base | — | https://api.openai.com/v1 | Trailing slash trimmed |
max_tokens | — | 512 | Clamped to 64–4096 |
organization | — | — | Sent as the OpenAI-Organization header |
organization
Only needed when the key belongs to several organizations and you want a
specific one billed and rate-limited. It is a plain identifier, not a secret,
so it sits in the file rather than behind an *_env indirection.
max_tokens is clamped, not rejected
Below 64 the classifier's JSON gets truncated mid-object and the reply fails to parse — which the pipeline reads as "no opinion", so you would lose the AI layer silently. Above 4096 a runaway reply costs money without saying anything more. Out-of-range values are clamped to the nearest bound rather than failing startup.
Choosing a model
| Want | Try |
|---|---|
| Cheapest sensible default | gpt-5-mini (the built-in default) |
| Better judgement on borderline messages | a full gpt-5-class model |
| Azure | your deployment name, not the underlying model name |
The dashboard's model picker calls GET /models and filters the account's
catalog down to models that can actually answer a chat completion — embedding,
whisper, tts and dall-e families are excluded, because offering them in a
picker only produces broken configs.
Set model explicitly. The default is chosen for cheap high-volume
classification, and inheriting it means the model changes under you when the
default does.
Azure OpenAI
Point api_base at the deployment URL:
[llm.providers.openai]
api_key_env = "AZURE_OPENAI_KEY"
api_base = "https://ACCOUNT.openai.azure.com/openai/deployments/DEPLOYMENT"
model = "DEPLOYMENT"
api-version and its own headerIf your Azure endpoint requires the api-version query parameter to survive
onto every request, or authenticates with the api-key header instead of
Authorization: Bearer, use the
OpenAI-compatible provider instead — it
preserves a query string on api_base and passes arbitrary headers through.
Cost and privacy
- Only ambiguous messages are sent. Clear passes and textbook fails are
decided locally. Under
sensitivity = "conservative"nothing is sent at all. - What is sent: the message text plus the team charter; for a rewrite, any resolved context items too.
- Replies are capped at
max_tokens(512 by default) on both operations. - Requests time out after 30 seconds and fall through to the next provider in the chain.
- If message content must not reach a vendor at all, use Ollama.
Errors and secrets
Failures carry the HTTP status and a capped excerpt of the response body —
never the API key, which lives only in the request header. The provider struct
deliberately does not derive Debug, because a derived Debug is how a secret
reaches a log line by accident.
A 401 is the key; a 404 on a valid key is usually a model the account cannot
reach; a 429 is what the fallback chain exists
for.