Ollama — local and private
This is the privacy and offline story. Ollama runs models on hardware you own, needs no API key, and talks to a host on your own network — so message text never leaves it.
Two things follow from that, and they are the two reasons to use it:
- It is the only provider you can point at conversations nobody is allowed to hand to a vendor.
- It still answers with the internet unplugged, which makes it the right last entry in a fallback chain.
Configuration
Everything is optional. A bare section is a complete configuration:
[llm]
provider = "ollama"
[llm.providers.ollama]
# every key below has a working default
More realistically, as the last rung behind a hosted primary:
[llm]
provider = "openrouter"
fallback = ["ollama"]
[llm.providers.openrouter]
api_key_env = "OPENROUTER_API_KEY"
model = "anthropic/claude-sonnet-5"
[llm.providers.ollama]
model = "llama3.1:8b"
api_base = "http://ollama.internal:11434/v1"
| Key | Required | Default | Notes |
|---|---|---|---|
model | — | llama3.1:8b | Any model you have pulled |
api_base | — | http://localhost:11434/v1 | Ollama's OpenAI-compatible surface, not its native /api/* one |
max_tokens | — | 512 | Cap on the reply |
api_key_env | — | — | Optional; only for a reverse proxy in front of Ollama that demands a bearer token |
/v1, not /apiOllama exposes two surfaces. This provider speaks the OpenAI-compatible one at
/v1. Pointing api_base at http://localhost:11434 or
http://localhost:11434/api will not work.
Why the timeout is 120 seconds
Every hosted provider times out after 30 seconds. This one waits 120:
Local inference on CPU is minutes-slow where a hosted provider is seconds-slow, so the hosted 30s cap would time out a working server.
That is the trade being made. A CPU-only box running an 8B model may take tens
of seconds per classification, and the nudge grace window is 150 seconds by
default — so a slow local model still fits, but not with much room. If you are
running Ollama as the primary provider on modest hardware, consider raising
[policy] nudge_delay_secs.
Choosing a model
| Hardware | Try |
|---|---|
| CPU only, modest box | a 3B–8B instruct model — llama3.2:3b, llama3.1:8b (the default) |
| A single consumer GPU | a 7B–14B instruct model, quantized |
| A real GPU box | a 30B+ instruct model, or run vLLM instead |
Two things matter more than raw size for this job:
- Instruction following. The classifier must return one strict JSON object. A model that chats around the answer fails the parse — which is read as "no opinion", so you quietly lose the AI layer. Prefer an instruct/chat-tuned model over a base model.
- Latency. This runs on every ambiguous message.
Pull it first:
ollama pull llama3.1:8b
ollama list
list_models is backed by the server's own catalog, so the dashboard picker
shows exactly the models that host has pulled.
Running it somewhere other than localhost
Ollama binds 127.0.0.1 by default. To reach it from another host:
OLLAMA_HOST=0.0.0.0:11434 ollama serve
[llm.providers.ollama]
api_base = "http://192.168.1.50:11434/v1"
Ollama has no authentication of its own — anyone who can reach the port can use the models. Keep it on a private network, or put a reverse proxy in front of it and use the optional bearer token:
[llm.providers.ollama]
api_base = "https://ollama.internal/v1"
api_key_env = "OLLAMA_PROXY_TOKEN"
The key is genuinely optional: absent means no Authorization header at all,
which is what a bare local server expects. When api_key_env is present it
must resolve — a named-but-unset variable is a real config error, not a
tolerable absence.
Same prompts, same verdicts
Only the endpoint, the optional auth header and the timeout are local concerns.
The prompts and every reply parse come from the shared crate, so a local
verdict is produced from exactly the same prompt as a hosted one. That is
what makes openrouter → ollama a safe chain rather than a personality change
halfway down.
Cost and privacy
| Cost per message | zero — your hardware |
| Message text leaves your network | never |
| Works offline | yes |
| Quality vs. a frontier model | lower, especially on borderline judgement |
If your reason for using Ollama is privacy, remember that
sensitivity = "conservative" sends nothing to any provider — the cheapest
private configuration is no AI layer at all. Ollama is what you use when you
want the ambiguous middle judged and the text kept in the building.