Skip to main content

Ollama — local and private

This is the privacy and offline story. Ollama runs models on hardware you own, needs no API key, and talks to a host on your own network — so message text never leaves it.

Two things follow from that, and they are the two reasons to use it:

  1. It is the only provider you can point at conversations nobody is allowed to hand to a vendor.
  2. It still answers with the internet unplugged, which makes it the right last entry in a fallback chain.

Configuration​

Everything is optional. A bare section is a complete configuration:

[llm]
provider = "ollama"

[llm.providers.ollama]
# every key below has a working default

More realistically, as the last rung behind a hosted primary:

[llm]
provider = "openrouter"
fallback = ["ollama"]

[llm.providers.openrouter]
api_key_env = "OPENROUTER_API_KEY"
model = "anthropic/claude-sonnet-5"

[llm.providers.ollama]
model = "llama3.1:8b"
api_base = "http://ollama.internal:11434/v1"
KeyRequiredDefaultNotes
model—llama3.1:8bAny model you have pulled
api_base—http://localhost:11434/v1Ollama's OpenAI-compatible surface, not its native /api/* one
max_tokens—512Cap on the reply
api_key_env——Optional; only for a reverse proxy in front of Ollama that demands a bearer token
warning
/v1, not /api

Ollama exposes two surfaces. This provider speaks the OpenAI-compatible one at /v1. Pointing api_base at http://localhost:11434 or http://localhost:11434/api will not work.

Why the timeout is 120 seconds​

Every hosted provider times out after 30 seconds. This one waits 120:

Local inference on CPU is minutes-slow where a hosted provider is seconds-slow, so the hosted 30s cap would time out a working server.

That is the trade being made. A CPU-only box running an 8B model may take tens of seconds per classification, and the nudge grace window is 150 seconds by default — so a slow local model still fits, but not with much room. If you are running Ollama as the primary provider on modest hardware, consider raising [policy] nudge_delay_secs.

Choosing a model​

HardwareTry
CPU only, modest boxa 3B–8B instruct model — llama3.2:3b, llama3.1:8b (the default)
A single consumer GPUa 7B–14B instruct model, quantized
A real GPU boxa 30B+ instruct model, or run vLLM instead

Two things matter more than raw size for this job:

  1. Instruction following. The classifier must return one strict JSON object. A model that chats around the answer fails the parse — which is read as "no opinion", so you quietly lose the AI layer. Prefer an instruct/chat-tuned model over a base model.
  2. Latency. This runs on every ambiguous message.

Pull it first:

ollama pull llama3.1:8b
ollama list

list_models is backed by the server's own catalog, so the dashboard picker shows exactly the models that host has pulled.

Running it somewhere other than localhost​

Ollama binds 127.0.0.1 by default. To reach it from another host:

OLLAMA_HOST=0.0.0.0:11434 ollama serve
[llm.providers.ollama]
api_base = "http://192.168.1.50:11434/v1"

Ollama has no authentication of its own — anyone who can reach the port can use the models. Keep it on a private network, or put a reverse proxy in front of it and use the optional bearer token:

[llm.providers.ollama]
api_base = "https://ollama.internal/v1"
api_key_env = "OLLAMA_PROXY_TOKEN"

The key is genuinely optional: absent means no Authorization header at all, which is what a bare local server expects. When api_key_env is present it must resolve — a named-but-unset variable is a real config error, not a tolerable absence.

Same prompts, same verdicts​

Only the endpoint, the optional auth header and the timeout are local concerns. The prompts and every reply parse come from the shared crate, so a local verdict is produced from exactly the same prompt as a hosted one. That is what makes openrouter → ollama a safe chain rather than a personality change halfway down.

Cost and privacy​

Cost per messagezero — your hardware
Message text leaves your networknever
Works offlineyes
Quality vs. a frontier modellower, especially on borderline judgement

If your reason for using Ollama is privacy, remember that sensitivity = "conservative" sends nothing to any provider — the cheapest private configuration is no AI layer at all. Ollama is what you use when you want the ambiguous middle judged and the text kept in the building.