Skip to main content

Grok

xAI's Grok, over its OpenAI-compatible chat-completions API at api.x.ai.

xAI speaks the same wire shape as OpenAI, so this provider is little more than an endpoint and an auth header: the request body, both prompts and every reply parse come from the shared crate.

Good for​

  • Teams that already have an xAI account and want to keep AI spend on one bill.
  • A drop-in second rung in a fallback chain — a different vendor with a different outage profile from OpenAI or Anthropic.

Configuration​

[llm]
provider = "grok"

[llm.providers.grok]
api_key_env = "XAI_API_KEY" # required
model = "grok-4" # set this explicitly
max_tokens = 512
# api_base = "https://api.x.ai/v1"
export XAI_API_KEY=xai-...
KeyRequiredDefaultNotes
api_key_env✅—Env var holding the key
model—grok-4Any chat model the account can reach
api_base—https://api.x.ai/v1Trailing slash trimmed
max_tokens—512Cap on the reply

Choosing a model​

xAI's line-up is small, which makes this easy: take the current general chat model unless you have a reason not to. grok-4 is the built-in default; a -mini or -fast variant is the right move if you are running high sensitivity on a busy workspace and want the classifier cheaper.

Set model explicitly — the default will move as xAI's line-up does, and you want that change to be yours.

list_models is backed by the account's catalog where xAI exposes one; if it returns nothing, type the id.

Why the prompts are not xAI-specific​

Every provider sends the same two prompts:

A provider that invented its own prompt would make the fallback chain change its mind about a message purely by moving down a rung, which is the one thing the chain must not do.

So a verdict from Grok is produced from exactly the same question as a verdict from OpenAI, and the only thing that varies is the model's judgement. That is what makes a chain like grok → openai → ollama safe to run.

The classifier asks for response_format: {"type":"json_object"}; the rewrite does not, because a rewrite is prose the author posts verbatim. Gateways that ignore the field degrade gracefully — the parser tolerates fenced and prose-wrapped JSON.

Cost and privacy​

  • Only ambiguous messages are sent, and only under balanced / high. Under conservative nothing is sent at all.
  • What is sent: the message text and the team charter; for a rewrite, resolved context items too.
  • Check xAI's current data-retention and training terms for your account tier before pointing this at real work conversations.
  • Requests time out after 30 seconds — hosted inference answers in seconds, and a longer wait is a stuck request. The pipeline would rather fall back than hold a nudge open.
  • If message content must not leave your network, use Ollama.

Errors and secrets​

Malformed replies are errors, never guesses; the pipeline reads an LLM error as "no opinion" and falls back to heuristics.

Errors carry the status and a capped body excerpt — never the key. The Authorization header value is built in one place, kept separate from the request so its shape is testable without the value ever reaching a log line.

A 401 is the key; a 404 on a valid key is a model id the account cannot reach; a 429 is what the fallback chain exists for.