Grok
xAI's Grok, over its OpenAI-compatible chat-completions API at api.x.ai.
xAI speaks the same wire shape as OpenAI, so this provider is little more than an endpoint and an auth header: the request body, both prompts and every reply parse come from the shared crate.
Good for
- Teams that already have an xAI account and want to keep AI spend on one bill.
- A drop-in second rung in a fallback chain — a different vendor with a different outage profile from OpenAI or Anthropic.
Configuration
[llm]
provider = "grok"
[llm.providers.grok]
api_key_env = "XAI_API_KEY" # required
model = "grok-4" # set this explicitly
max_tokens = 512
# api_base = "https://api.x.ai/v1"
export XAI_API_KEY=xai-...
| Key | Required | Default | Notes |
|---|---|---|---|
api_key_env | ✅ | — | Env var holding the key |
model | — | grok-4 | Any chat model the account can reach |
api_base | — | https://api.x.ai/v1 | Trailing slash trimmed |
max_tokens | — | 512 | Cap on the reply |
Choosing a model
xAI's line-up is small, which makes this easy: take the current general chat
model unless you have a reason not to. grok-4 is the built-in default; a
-mini or -fast variant is the right move if you are running high
sensitivity on a busy workspace and want the classifier cheaper.
Set model explicitly — the default will move as xAI's line-up does, and you
want that change to be yours.
list_models is backed by the account's catalog where xAI exposes one; if it
returns nothing, type the id.
Why the prompts are not xAI-specific
Every provider sends the same two prompts:
A provider that invented its own prompt would make the fallback chain change its mind about a message purely by moving down a rung, which is the one thing the chain must not do.
So a verdict from Grok is produced from exactly the same question as a verdict
from OpenAI, and the only thing that varies is the model's judgement. That is
what makes a chain like grok → openai → ollama safe to run.
The classifier asks for response_format: {"type":"json_object"}; the rewrite
does not, because a rewrite is prose the author posts verbatim. Gateways that
ignore the field degrade gracefully — the parser tolerates fenced and
prose-wrapped JSON.
Cost and privacy
- Only ambiguous messages are sent, and only under
balanced/high. Underconservativenothing is sent at all. - What is sent: the message text and the team charter; for a rewrite, resolved context items too.
- Check xAI's current data-retention and training terms for your account tier before pointing this at real work conversations.
- Requests time out after 30 seconds — hosted inference answers in seconds, and a longer wait is a stuck request. The pipeline would rather fall back than hold a nudge open.
- If message content must not leave your network, use Ollama.
Errors and secrets
Malformed replies are errors, never guesses; the pipeline reads an LLM error as "no opinion" and falls back to heuristics.
Errors carry the status and a capped body excerpt — never the key. The
Authorization header value is built in one place, kept separate from the
request so its shape is testable without the value ever reaching a log line.
A 401 is the key; a 404 on a valid key is a model id the account cannot reach; a 429 is what the fallback chain exists for.