Operations
Everything you need to know whether Ever Async is alive, working, and worth keeping installed.
| Surface | What it answers |
|---|---|
GET /healthz | Is the process up? |
GET /metrics | What has it been doing? |
GET /api/v1/status | What is registered? |
GET /api/v1/digest | Was it worth it this week? |
GET /api/v1/auth/mode | Is anything actually being enforced? |
everasync doctor | Are the credentials still good — and is the API exposed? |
Every curl on this page is written for the default [auth] mode = "local_trusted", where no credential is required. Under
mode = "authenticated" add -H "Authorization: Bearer …" to all of them except
/healthz and /api/v1/auth/mode — see
Security → which role reaches which endpoint.
Health check
curl -s http://localhost:8100/healthz # → ok
/healthz returns the literal string ok and touches nothing — no database,
no plugin, no network. That is what makes it safe as a Kubernetes liveness and
readiness probe.
Pointing a probe at / or an /api/v1/... route makes a slow dependency
restart the pod, which makes the dependency slower. Use /healthz.
Metrics
curl -s http://localhost:8100/metrics
Prometheus text exposition format
(text/plain; version=0.0.4; charset=utf-8), served at the root path by
convention so a stock scrape config finds it with no path override. It is
outside the CORS layer — a scraper is not a browser.
Every metric is a counter. Series are sorted, so scrapes and diffs are stable.
Counters are created the first time something increments them, so a server that
has not seen a webhook yet answers 200 with an empty body. That is a valid
Prometheus scrape, not a broken endpoint — check the status code and the
content-type, not the length. Post one message in a watched channel and
ever_async_events_total appears.
The full counter list
| Counter | Labels | Meaning |
|---|---|---|
ever_async_events_total | channel | Messages observed |
ever_async_events_duplicate_total | channel | Webhook redeliveries suppressed by idempotency |
ever_async_nudges_total | channel | Private nudges delivered |
ever_async_nudges_cancelled_total | — | Nudges dropped because the author self-corrected |
ever_async_nudges_claim_lost_total | — | Nudge timers that fired after another replica had claimed the nudge. ~0 at one replica; grows with replicas - 1 per nudge above that |
ever_async_nudges_stale_dropped_total | — | Pending nudges dropped at startup for being overdue beyond [nudges] max_overdue_secs. 0 after an ordinary restart; non-zero counts the messages an outage cost you |
ever_async_rewrites_accepted_total | channel | Suggested rewrites the author chose to post |
ever_async_context_cards_total | channel | Context cards attached to threads |
ever_async_connector_calls_total | connector, outcome | Connector resolutions. outcome ∈ ok · error · timeout |
ever_async_llm_calls_total | provider, operation, outcome | LLM calls. operation ∈ classify · suggest_rewrite; outcome ∈ ok · error |
ever_async_ingress_rejected_total | reason | Ingress rejected. reason ∈ verification · error |
ever_async_actions_total | channel, kind | Outbound actions. kind ∈ ephemeral · ephemeral_buttons · thread_reply · post · direct_message · replace |
Sample scrape:
# HELP ever_async_events_total Messages observed, by channel.
# TYPE ever_async_events_total counter
ever_async_events_total{channel="slack"} 1842
# TYPE ever_async_nudges_total counter
ever_async_nudges_total{channel="slack"} 37
# TYPE ever_async_nudges_cancelled_total counter
ever_async_nudges_cancelled_total 61
# TYPE ever_async_connector_calls_total counter
ever_async_connector_calls_total{connector="github",outcome="ok"} 214
ever_async_connector_calls_total{connector="jira",outcome="timeout"} 3
There is no persistence behind the metric registry. Use increase() /
rate() in your queries rather than raw values, and read the
digest — which is computed from storage — for anything that must
survive a restart.
/metrics is not on the public allowlistCounter names and label values carry no message content — but they name every
configured channel, count every conversation, and show which connectors are
erroring and which LLM provider is in use. So under
[auth] mode = "authenticated" a scrape needs a viewer credential like any
other caller. In the default local_trusted mode it needs none, which is why the
curl above works out of the box.
Mint a viewer token and give it to Prometheus:
scrape_configs:
- job_name: everasync
authorization:
type: Bearer
credentials_file: /etc/prometheus/everasync-token
static_configs:
- targets: ['everasync:8100']
/metrics sits outside the CORS layer (a scraper is not a browser), which is
a different thing from sitting outside the auth layer.
What to actually watch
Is it working at all?
rate(ever_async_events_total[5m]) # traffic is arriving
increase(ever_async_ingress_rejected_total{reason="verification"}[1h])
Any sustained signature-rejection rate means a wrong signing secret, a proxy rewriting the body, or host clock skew — see Self-hosting → Reverse proxy.
Is it being useful, or annoying? The three numbers that matter:
increase(ever_async_nudges_total[7d])
increase(ever_async_nudges_cancelled_total[7d])
increase(ever_async_rewrites_accepted_total[7d])
- A high cancel ratio is good. People are self-correcting inside the grace window; the nudge was never needed. That is the behaviour change you wanted.
rewrites_accepted / nudgesis the honest quality signal. If nobody ever presses Post this, the suggestions are not worth a click — change the model or drop the sensitivity.- A rising
nudges_totalwith a flatrewrites_accepted_totalis the shape of a tool people are learning to ignore.
Are the connectors healthy?
sum by (connector) (rate(ever_async_connector_calls_total{outcome!="ok"}[15m]))
Timeouts are the interesting one: the per-connector budget is 6 seconds, and a connector that times out is silently contributing nothing to every message.
Is the AI layer earning its keep?
sum by (provider, outcome) (rate(ever_async_llm_calls_total[15m]))
Errors on the primary with successes on a fallback is the chain doing its job. Errors on every provider means the AI layer is effectively off and only heuristics are running — which the product survives, but you should know about it.
The digest
The value-proof loop: what did Ever Async actually do this week?
curl -s 'http://localhost:8100/api/v1/digest?days=7'
{ "events_seen": 1842, "nudges_sent": 37, "context_items_attached": 96 }
days defaults to 7 and is clamped to 1–90. Unlike /metrics, this is
computed from storage, so it survives restarts.
Same thing from inside a channel, replied privately:
/async digest
📊 Ever Async — last 7 days
• Messages observed: 1842
• Private nudges sent: 37
• Context links attached: 96
Delivery
[digest]
enabled = true
interval_hours = 168 # weekly
days = 7
channel = "slack" # DM through this channel plugin
recipients = ["U0123ABC"]
notify = false # also hand it to every notify provider (Novu, …)
Send one right now, without waiting for a schedule:
curl -X POST http://localhost:8100/api/v1/digest/send
This uses the [digest] settings for the window, the channel and the recipient
list.
everasync serve calls Pipeline::start_background once — from
run_with_sso, which run delegates to, so both entry points get it exactly
once. It does two things, and they have very different multi-replica stories.
Pending nudges now survive a restart. They were always persisted; nothing
read them back. Boot re-arms them, and that is safe at any replica count:
load_pending_nudges is global, so every replica restores every nudge and arms
its own timer, and claim_pending_nudge elects the single one that delivers.
Nudges that came due while the process was down are re-armed too — but only up
to [nudges] max_overdue_secs (default 15 min). Anything staler is dropped at
boot and counted in ever_async_nudges_stale_dropped_total; see
Restoring the queue below.
The digest loop is the part that is still per-process. It starts only when
[digest] enabled = true (the default is false), and it is a plain
tokio::time::interval with no claim behind it, so N replicas deliver N
digests per interval. Either keep it disabled while replicas > 1 and drive
POST /api/v1/digest/send from a single cron job or Kubernetes CronJob — a
more observable schedule anyway — or leave it on and accept a digest per pod.
It also delivers the default tenant's counters only.
Restoring the queue after a restart
A nudge waits out a grace period ([policy] nudge_delay_secs, default 150s) on
an in-process timer. If the process stops during that window the row outlives it
and the timer does not, so boot has to decide what to do with a nudge whose
waiting is already over.
Firing all of them immediately is the wrong answer. After a 30-second rolling
deploy it is right — the message is minutes old, the thread is live. After a
two-hour outage it means a pile of nudges landing at once, every one about a
message nobody is reading any more. [nudges] max_overdue_secs is where that
line is drawn:
| Restored nudge | What boot does |
|---|---|
| Due in the future | Re-armed for the time it has left |
Overdue by ≤ max_overdue_secs | Armed immediately — still runs the self-correction check and the claim |
| Overdue by more | Claimed and dropped. ever_async_nudges_stale_dropped_total +1 |
The drop takes the same claim a delivery takes, so the row is consumed once fleet-wide and the counter moves once however many replicas booted — and a dropped nudge is never also delivered somewhere else.
increase(ever_async_nudges_stale_dropped_total[1d])
Zero is the normal reading. Anything else says a restart took longer than
max_overdue_secs and names how many messages went unanswered because of it —
after an incident, this is the number that says what the outage cost in nudges.
everasync doctor
everasync doctor
everasync doctor --config /etc/everasync.toml
Loads the config, assembles every plugin exactly the way serve does, then
health-checks each channel and connector concurrently. Errors are collected as
text so one bad plugin never aborts the sweep.
✓ channel slack
✓ connector github
✗ connector jira jira /rest/api/3/myself returned 404
· storage sqlite everasync.db — one writer, so one replica
· llm openrouter anthropic/claude-sonnet-5
· llm ollama llama3.1:8b
! auth disabled WARNING: no credential is required and the server is bound to 0.0.0.0:8100 …
Four markers:
| Marker | Meaning |
|---|---|
✓ | health check passed |
✗ | health check failed — counts toward the non-zero exit |
! | a warning. Exactly one thing raises it today: auth is off and the bind is not loopback |
· | informational, not a verdict |
Two rows are always present, whatever is configured:
storage— the driver actually opened, after defaulting, plus where the data lives. Assembly already opened the file or connected the pool, so a broken backend never reaches this line; and for Postgres the detail is the name of the environment variable, never the connection string.auth— printed last, so the answer to "is this thing exposed?" is the line left on your screen. It names the provider enforcing credentials, or saysdisabledand explains what that means for the bind in use.
The LLM chain is listed in routing order with · because LlmProvider has no
health check — probing one would spend tokens on every doctor run. It is
still the fastest way to see the whole chain and the model each rung is pointed
at, since /api/v1/status reports the router as a single provider.
Exits non-zero when any check fails, which is what makes it usable in CI and after every credential rotation. With nothing configured it says so rather than reporting a vacuous success.
What each check actually calls:
| Plugin | Health check |
|---|---|
| Slack | auth.test |
| GitHub | an authenticated API call |
| Jira | GET /rest/api/3/myself |
| Ever Gauzy | an authenticated API call with the tenant header |
doctor does not exercise the LLM chain or spend tokens. To verify a
provider, set sensitivity = "balanced", post an ambiguous message, and watch
ever_async_llm_calls_total.
Inventory
curl -s http://localhost:8100/api/v1/status | jq
{"plugins":[
{"id":"slack","name":"Slack","version":"0.1.0","kind":"channel"},
{"id":"github","name":"GitHub","version":"0.1.0","kind":"connector"},
{"id":"openrouter","name":"OpenRouter (+1 fallback)","version":"0.1.0",
"kind":"llm","model":"anthropic/claude-sonnet-5"}
]}
An empty plugins array means no section matched — the most common cause of
"it does nothing". The LLM entry reports the router, so a chain shows as
its primary plus a fallback count; everasync doctor lists the rungs
individually.
The rest of the API
The Least role column is what a caller needs under
[auth] mode = "authenticated". In the default local_trusted mode every caller
is the local board principal, so every row passes and nothing 401s or 403s.
| Route | Method | Least role | Purpose |
|---|---|---|---|
/healthz | GET | public | Liveness — returns ok |
/ingress/{channel}[/…] | POST | public — signature-verified instead | Webhook intake |
/metrics | GET | viewer | Prometheus counters |
/api/v1/status | GET | viewer | Registered plugins |
/api/v1/charter | GET · PUT | viewer · operator | Read / replace the rules every nudge cites |
/api/v1/digest | GET | viewer | Counters (?days=N, 1–90) |
/api/v1/digest/send | POST | board | Deliver one digest now |
/api/v1/llm/models | GET | viewer | Models the primary provider can be pointed at |
/api/v1/policies/{channel}/{conversation} | GET · PUT | viewer · operator | Per-conversation policy |
/api/v1/auth/mode | GET | public | What the server is enforcing, and which sign-in buttons to render |
/api/v1/auth/me | GET | any authenticated caller | The principal this request resolved to |
/api/v1/auth/login | POST | public | Password → session. Mounted only when the provider issues sessions (local) |
/api/v1/auth/oauth/{provider}/start | GET | public | Begin browser SSO. Mounted only when [auth.providers.oauth] is configured |
/api/v1/auth/oauth/{provider}/callback | GET | public | Where the identity provider sends the browser back. Same condition |
/api/v1/auth/switch-tenant | POST | authenticated, and a member of the target tenant | Re-issue the session bound to another tenant. Same condition |
Routes marked mounted only do not exist at all on a deployment that did not configure them — there is nothing to probe, and an unconfigured provider id 404s exactly like a route that was never mounted.
/api/v1 is writable, and unauthenticated by default[auth] mode defaults to local_trusted, which asks for no credential at all,
and /api/v1 carries permissive CORS. PUT /api/v1/charter rewrites the rules
every nudge cites, PUT /api/v1/policies/… can silence the bot in a channel, and
POST /api/v1/digest/send messages real people.
Do not expose this to the internet without turning on
[auth]. Ingress is different — it is protected by each
platform's request signature, which is why it stays anonymous.
Logs
RUST_LOG wins when set; otherwise Ever Async's own crates run at debug and
the ecosystem at info.
RUST_LOG=info everasync serve
RUST_LOG=ever_async_core::pipeline=debug,info everasync serve
Warnings worth alerting on:
| Log line | Meaning |
|---|---|
connector resolve timed out | a connector is over its 6s budget on every message |
llm provider failed; trying next | fine occasionally; constant means the primary is down |
nudge delivery failed | the channel token lost a scope, or the bot was removed |
could not persist pending nudge | storage problem. The timer is armed anyway, but delivery goes through a claim on that row — so if the write really did not land, the nudge is dropped rather than sent unclaimed |
dropped pending nudges that were overdue beyond the cutoff | this process was down for longer than [nudges] max_overdue_secs; the count is how many nudges that cost |
duplicate delivery ignored | normal. Platforms retry; this is dedup working |