Skip to main content

Operations

Everything you need to know whether Ever Async is alive, working, and worth keeping installed.

SurfaceWhat it answers
GET /healthzIs the process up?
GET /metricsWhat has it been doing?
GET /api/v1/statusWhat is registered?
GET /api/v1/digestWas it worth it this week?
GET /api/v1/auth/modeIs anything actually being enforced?
everasync doctorAre the credentials still good — and is the API exposed?

Every curl on this page is written for the default [auth] mode = "local_trusted", where no credential is required. Under mode = "authenticated" add -H "Authorization: Bearer …" to all of them except /healthz and /api/v1/auth/mode — see Security → which role reaches which endpoint.

Health check​

curl -s http://localhost:8100/healthz     # → ok

/healthz returns the literal string ok and touches nothing — no database, no plugin, no network. That is what makes it safe as a Kubernetes liveness and readiness probe.

Never probe an expensive route

Pointing a probe at / or an /api/v1/... route makes a slow dependency restart the pod, which makes the dependency slower. Use /healthz.

Metrics​

curl -s http://localhost:8100/metrics

Prometheus text exposition format (text/plain; version=0.0.4; charset=utf-8), served at the root path by convention so a stock scrape config finds it with no path override. It is outside the CORS layer — a scraper is not a browser.

Every metric is a counter. Series are sorted, so scrapes and diffs are stable.

A fresh install scrapes empty

Counters are created the first time something increments them, so a server that has not seen a webhook yet answers 200 with an empty body. That is a valid Prometheus scrape, not a broken endpoint — check the status code and the content-type, not the length. Post one message in a watched channel and ever_async_events_total appears.

The full counter list​

CounterLabelsMeaning
ever_async_events_totalchannelMessages observed
ever_async_events_duplicate_totalchannelWebhook redeliveries suppressed by idempotency
ever_async_nudges_totalchannelPrivate nudges delivered
ever_async_nudges_cancelled_total—Nudges dropped because the author self-corrected
ever_async_nudges_claim_lost_total—Nudge timers that fired after another replica had claimed the nudge. ~0 at one replica; grows with replicas - 1 per nudge above that
ever_async_nudges_stale_dropped_total—Pending nudges dropped at startup for being overdue beyond [nudges] max_overdue_secs. 0 after an ordinary restart; non-zero counts the messages an outage cost you
ever_async_rewrites_accepted_totalchannelSuggested rewrites the author chose to post
ever_async_context_cards_totalchannelContext cards attached to threads
ever_async_connector_calls_totalconnector, outcomeConnector resolutions. outcome ∈ ok · error · timeout
ever_async_llm_calls_totalprovider, operation, outcomeLLM calls. operation ∈ classify · suggest_rewrite; outcome ∈ ok · error
ever_async_ingress_rejected_totalreasonIngress rejected. reason ∈ verification · error
ever_async_actions_totalchannel, kindOutbound actions. kind ∈ ephemeral · ephemeral_buttons · thread_reply · post · direct_message · replace

Sample scrape:

# HELP ever_async_events_total Messages observed, by channel.
# TYPE ever_async_events_total counter
ever_async_events_total{channel="slack"} 1842
# TYPE ever_async_nudges_total counter
ever_async_nudges_total{channel="slack"} 37
# TYPE ever_async_nudges_cancelled_total counter
ever_async_nudges_cancelled_total 61
# TYPE ever_async_connector_calls_total counter
ever_async_connector_calls_total{connector="github",outcome="ok"} 214
ever_async_connector_calls_total{connector="jira",outcome="timeout"} 3
Counters are in-memory and reset on restart

There is no persistence behind the metric registry. Use increase() / rate() in your queries rather than raw values, and read the digest — which is computed from storage — for anything that must survive a restart.

warning
/metrics is not on the public allowlist

Counter names and label values carry no message content — but they name every configured channel, count every conversation, and show which connectors are erroring and which LLM provider is in use. So under [auth] mode = "authenticated" a scrape needs a viewer credential like any other caller. In the default local_trusted mode it needs none, which is why the curl above works out of the box.

Mint a viewer token and give it to Prometheus:

scrape_configs:
- job_name: everasync
authorization:
type: Bearer
credentials_file: /etc/prometheus/everasync-token
static_configs:
- targets: ['everasync:8100']

/metrics sits outside the CORS layer (a scraper is not a browser), which is a different thing from sitting outside the auth layer.

What to actually watch​

Is it working at all?

rate(ever_async_events_total[5m])          # traffic is arriving
increase(ever_async_ingress_rejected_total{reason="verification"}[1h])

Any sustained signature-rejection rate means a wrong signing secret, a proxy rewriting the body, or host clock skew — see Self-hosting → Reverse proxy.

Is it being useful, or annoying? The three numbers that matter:

increase(ever_async_nudges_total[7d])
increase(ever_async_nudges_cancelled_total[7d])
increase(ever_async_rewrites_accepted_total[7d])
  • A high cancel ratio is good. People are self-correcting inside the grace window; the nudge was never needed. That is the behaviour change you wanted.
  • rewrites_accepted / nudges is the honest quality signal. If nobody ever presses Post this, the suggestions are not worth a click — change the model or drop the sensitivity.
  • A rising nudges_total with a flat rewrites_accepted_total is the shape of a tool people are learning to ignore.

Are the connectors healthy?

sum by (connector) (rate(ever_async_connector_calls_total{outcome!="ok"}[15m]))

Timeouts are the interesting one: the per-connector budget is 6 seconds, and a connector that times out is silently contributing nothing to every message.

Is the AI layer earning its keep?

sum by (provider, outcome) (rate(ever_async_llm_calls_total[15m]))

Errors on the primary with successes on a fallback is the chain doing its job. Errors on every provider means the AI layer is effectively off and only heuristics are running — which the product survives, but you should know about it.

The digest​

The value-proof loop: what did Ever Async actually do this week?

curl -s 'http://localhost:8100/api/v1/digest?days=7'
{ "events_seen": 1842, "nudges_sent": 37, "context_items_attached": 96 }

days defaults to 7 and is clamped to 1–90. Unlike /metrics, this is computed from storage, so it survives restarts.

Same thing from inside a channel, replied privately:

/async digest
📊 Ever Async — last 7 days
• Messages observed: 1842
• Private nudges sent: 37
• Context links attached: 96

Delivery​

[digest]
enabled = true
interval_hours = 168 # weekly
days = 7
channel = "slack" # DM through this channel plugin
recipients = ["U0123ABC"]
notify = false # also hand it to every notify provider (Novu, …)

Send one right now, without waiting for a schedule:

curl -X POST http://localhost:8100/api/v1/digest/send

This uses the [digest] settings for the window, the channel and the recipient list.

The scheduled loop starts, and it is still the per-process piece

everasync serve calls Pipeline::start_background once — from run_with_sso, which run delegates to, so both entry points get it exactly once. It does two things, and they have very different multi-replica stories.

Pending nudges now survive a restart. They were always persisted; nothing read them back. Boot re-arms them, and that is safe at any replica count: load_pending_nudges is global, so every replica restores every nudge and arms its own timer, and claim_pending_nudge elects the single one that delivers. Nudges that came due while the process was down are re-armed too — but only up to [nudges] max_overdue_secs (default 15 min). Anything staler is dropped at boot and counted in ever_async_nudges_stale_dropped_total; see Restoring the queue below.

The digest loop is the part that is still per-process. It starts only when [digest] enabled = true (the default is false), and it is a plain tokio::time::interval with no claim behind it, so N replicas deliver N digests per interval. Either keep it disabled while replicas > 1 and drive POST /api/v1/digest/send from a single cron job or Kubernetes CronJob — a more observable schedule anyway — or leave it on and accept a digest per pod. It also delivers the default tenant's counters only.

Restoring the queue after a restart​

A nudge waits out a grace period ([policy] nudge_delay_secs, default 150s) on an in-process timer. If the process stops during that window the row outlives it and the timer does not, so boot has to decide what to do with a nudge whose waiting is already over.

Firing all of them immediately is the wrong answer. After a 30-second rolling deploy it is right — the message is minutes old, the thread is live. After a two-hour outage it means a pile of nudges landing at once, every one about a message nobody is reading any more. [nudges] max_overdue_secs is where that line is drawn:

Restored nudgeWhat boot does
Due in the futureRe-armed for the time it has left
Overdue by ≤ max_overdue_secsArmed immediately — still runs the self-correction check and the claim
Overdue by moreClaimed and dropped. ever_async_nudges_stale_dropped_total +1

The drop takes the same claim a delivery takes, so the row is consumed once fleet-wide and the counter moves once however many replicas booted — and a dropped nudge is never also delivered somewhere else.

increase(ever_async_nudges_stale_dropped_total[1d])

Zero is the normal reading. Anything else says a restart took longer than max_overdue_secs and names how many messages went unanswered because of it — after an incident, this is the number that says what the outage cost in nudges.

everasync doctor​

everasync doctor
everasync doctor --config /etc/everasync.toml

Loads the config, assembles every plugin exactly the way serve does, then health-checks each channel and connector concurrently. Errors are collected as text so one bad plugin never aborts the sweep.

✓ channel   slack
✓ connector github
✗ connector jira jira /rest/api/3/myself returned 404
· storage sqlite everasync.db — one writer, so one replica
· llm openrouter anthropic/claude-sonnet-5
· llm ollama llama3.1:8b
! auth disabled WARNING: no credential is required and the server is bound to 0.0.0.0:8100 …

Four markers:

MarkerMeaning
✓health check passed
✗health check failed — counts toward the non-zero exit
!a warning. Exactly one thing raises it today: auth is off and the bind is not loopback
·informational, not a verdict

Two rows are always present, whatever is configured:

  • storage — the driver actually opened, after defaulting, plus where the data lives. Assembly already opened the file or connected the pool, so a broken backend never reaches this line; and for Postgres the detail is the name of the environment variable, never the connection string.
  • auth — printed last, so the answer to "is this thing exposed?" is the line left on your screen. It names the provider enforcing credentials, or says disabled and explains what that means for the bind in use.

The LLM chain is listed in routing order with · because LlmProvider has no health check — probing one would spend tokens on every doctor run. It is still the fastest way to see the whole chain and the model each rung is pointed at, since /api/v1/status reports the router as a single provider.

Exits non-zero when any check fails, which is what makes it usable in CI and after every credential rotation. With nothing configured it says so rather than reporting a vacuous success.

What each check actually calls:

PluginHealth check
Slackauth.test
GitHuban authenticated API call
JiraGET /rest/api/3/myself
Ever Gauzyan authenticated API call with the tenant header

doctor does not exercise the LLM chain or spend tokens. To verify a provider, set sensitivity = "balanced", post an ambiguous message, and watch ever_async_llm_calls_total.

Inventory​

curl -s http://localhost:8100/api/v1/status | jq
{"plugins":[
{"id":"slack","name":"Slack","version":"0.1.0","kind":"channel"},
{"id":"github","name":"GitHub","version":"0.1.0","kind":"connector"},
{"id":"openrouter","name":"OpenRouter (+1 fallback)","version":"0.1.0",
"kind":"llm","model":"anthropic/claude-sonnet-5"}
]}

An empty plugins array means no section matched — the most common cause of "it does nothing". The LLM entry reports the router, so a chain shows as its primary plus a fallback count; everasync doctor lists the rungs individually.

The rest of the API​

The Least role column is what a caller needs under [auth] mode = "authenticated". In the default local_trusted mode every caller is the local board principal, so every row passes and nothing 401s or 403s.

RouteMethodLeast rolePurpose
/healthzGETpublicLiveness — returns ok
/ingress/{channel}[/…]POSTpublic — signature-verified insteadWebhook intake
/metricsGETviewerPrometheus counters
/api/v1/statusGETviewerRegistered plugins
/api/v1/charterGET · PUTviewer · operatorRead / replace the rules every nudge cites
/api/v1/digestGETviewerCounters (?days=N, 1–90)
/api/v1/digest/sendPOSTboardDeliver one digest now
/api/v1/llm/modelsGETviewerModels the primary provider can be pointed at
/api/v1/policies/{channel}/{conversation}GET · PUTviewer · operatorPer-conversation policy
/api/v1/auth/modeGETpublicWhat the server is enforcing, and which sign-in buttons to render
/api/v1/auth/meGETany authenticated callerThe principal this request resolved to
/api/v1/auth/loginPOSTpublicPassword → session. Mounted only when the provider issues sessions (local)
/api/v1/auth/oauth/{provider}/startGETpublicBegin browser SSO. Mounted only when [auth.providers.oauth] is configured
/api/v1/auth/oauth/{provider}/callbackGETpublicWhere the identity provider sends the browser back. Same condition
/api/v1/auth/switch-tenantPOSTauthenticated, and a member of the target tenantRe-issue the session bound to another tenant. Same condition

Routes marked mounted only do not exist at all on a deployment that did not configure them — there is nothing to probe, and an unconfigured provider id 404s exactly like a route that was never mounted.

warning
/api/v1 is writable, and unauthenticated by default

[auth] mode defaults to local_trusted, which asks for no credential at all, and /api/v1 carries permissive CORS. PUT /api/v1/charter rewrites the rules every nudge cites, PUT /api/v1/policies/… can silence the bot in a channel, and POST /api/v1/digest/send messages real people.

Do not expose this to the internet without turning on [auth]. Ingress is different — it is protected by each platform's request signature, which is why it stays anonymous.

Logs​

RUST_LOG wins when set; otherwise Ever Async's own crates run at debug and the ecosystem at info.

RUST_LOG=info everasync serve
RUST_LOG=ever_async_core::pipeline=debug,info everasync serve

Warnings worth alerting on:

Log lineMeaning
connector resolve timed outa connector is over its 6s budget on every message
llm provider failed; trying nextfine occasionally; constant means the primary is down
nudge delivery failedthe channel token lost a scope, or the bot was removed
could not persist pending nudgestorage problem. The timer is armed anyway, but delivery goes through a claim on that row — so if the write really did not land, the nudge is dropped rather than sent unclaimed
dropped pending nudges that were overdue beyond the cutoffthis process was down for longer than [nudges] max_overdue_secs; the count is how many nudges that cost
duplicate delivery ignorednormal. Platforms retry; this is dedup working