The pipeline: classify → enrich → decide → act
Every message follows the same four stages. The interesting property is what each stage does when it is unsure: it hands the question to the next layer rather than guessing, and if no layer is confident, nothing happens at all.
0. Ingress — verify, then normalize
POST /ingress/{channel} (and /ingress/{channel}/{rest} for platforms that
suffix their routes) is the only inbound door. The HTTP layer does almost
nothing: it lower-cases the header names, keeps the raw body, and hands an
IngressRequest to the channel plugin named in the path.
The plugin verifies the platform's signature before parsing anything. A bad
signature is a hard Verification error → HTTP 401, counted in
ever_async_ingress_rejected_total{reason="verification"}. It is never a
warning and never falls through to processing.
What comes back is one of six outcomes — Events, Challenge, Command,
Interaction, Joined, or Ignored. Everything except Challenge and
Command is acknowledged immediately and processed on a spawned task:
webhooks must never wait on an LLM, because every platform retries a
delivery that looks slow.
Deduplication is the first thing that happens
Because platforms retry, the very first step of per-event processing claims the
message key ({channel}:{conversation}:{message_id}) in storage. A redelivery
loses the claim, increments ever_async_events_duplicate_total, and stops.
Without this, a slow ack would produce two nudges for one message — the
single worst failure mode this product has.
1. Classify — deterministic first
heuristic_classify runs first. It is pure regex over the message text: no
network, no tokens, microseconds. Three rules fire:
| Rule | Fires on | Verdict | Confidence |
|---|---|---|---|
| Naked greeting / contentless help plea | Hi team!, can I get some help? (≤ 12 words, no links) | LowContext, missing Details | 0.95 |
| Artifact without a link | a vague reference ("the PR", "my last task") and an action request (review, merge, deploy…) and no links | LowContext, missing Link | 0.90 |
| Clearly fine | has links, or ≥ 30 words | Clear | 0.90 |
Anything else is the ambiguous middle — short-ish, no links, no strong
pattern. The heuristics return Clear with confidence 0.55 and, crucially,
certain: false.
Only the ambiguous middle reaches an AI provider, and only when the conversation's sensitivity allows it:
| Sensitivity | LLM consulted? | Acts on an LLM verdict at |
|---|---|---|
conservative (default) | no | — heuristics only |
balanced | yes | confidence ≥ 0.80 |
high | yes | confidence ≥ 0.60 |
If the LLM call fails, the failure is logged and the heuristic verdict stands. An LLM error is read as "no opinion", never as "flag it".
2. Enrich — ask every connector at once
References extracted during classification ("the PR", EVER-123, PR #482)
go to every registered connector concurrently, each with a 6-second
timeout. A connector that is slow, broken or simply has nothing to say
contributes an empty list; it never blocks the others and never fails the
message. Outcomes are counted per connector in
ever_async_connector_calls_total{connector,outcome} with outcome one of
ok, error, timeout.
When [policy] unfurl_links is on (the default), the links the author
already posted are added to the same query. The platform's own preview shows
you a title; a connector shows you whether the PR is approved or the ticket is
blocked. Connectors claim URLs on hosts they own and ignore the rest — the core
never knows which host belongs to which tool.
Results are sorted by confidence and truncated to the top three. A context card is a nudge toward the answer, not a search-results page.
3. Decide — policy gates, in order
- Is the conversation enabled?
[policy] enabled, overridable per conversation with/async off. - Did anything resolve confidently? Items at or above
[policy] min_confidence(default0.7) count as strong. - Are we inside quiet hours? If
[policy.quiet_hours]is set and now is inside the window, no nudge is scheduled. Quiet hours suppress the nudge, not the context card.
4. Act — two very different outcomes
Strong context exists → a public thread reply. The card lists each item's
title, status and URL. No one is corrected; the message simply became
actionable. This path replaces the nudge — there is nothing left to ask the
author for. Counted in ever_async_context_cards_total.
Nothing resolved, and the message is low-context → a private nudge, later.
A PendingNudge is persisted and armed on an in-process timer for
[policy] nudge_delay_secs (default 150). When it fires, two checks run, in this
order:
- Self-correction. Storage is re-checked for an author follow-up; if the
author said anything in the thread meanwhile, the nudge is dropped and
ever_async_nudges_cancelled_totalgoes up. - The claim.
claim_pending_nudgedeletes the row and reports whether this process was the one that took it. Exactly one caller in the fleet is told yes.
Only then does the author — and nobody else — get an ephemeral message with the reasons, an optional AI-suggested rewrite, and a Post this button.
The order is deliberate. The timer map is per process, so with several replicas
every one of them arms a timer for the same nudge and all of them fire; the claim
is what makes that harmless. Taking it last, immediately before the send, keeps
the irreversible step (the row is deleted and nothing re-creates it) as close as
possible to the step it authorises — and keeps "the author already fixed it" a
decision about the message rather than about who won a race. A lost claim is the
ordinary multi-replica outcome, logged at debug and counted in
ever_async_nudges_claim_lost_total.
A restart does not drop the queue
The timer lives in the process; the PendingNudge lives in storage. When
everasync serve starts it re-arms everything still pending — one call to
Pipeline::start_background, from run_with_sso, which run delegates to.
load_pending_nudges is global, so every replica restores every nudge; all
those timers fire and the claim above elects the one that delivers. That is also
what keeps the restore robust: no nudge is tied to a replica that may not come
back.
The one judgement boot has to make is about nudges that came due while the process was down, because their grace period is already over:
| Restored nudge | Boot does |
|---|---|
| Due in the future | Re-arms it for the time it has left |
Overdue by ≤ [nudges] max_overdue_secs (default 15 min) | Arms it now — self-correction check and claim still run |
| Overdue by more | Claims and drops it; ever_async_nudges_stale_dropped_total +1 |
Firing every overdue nudge on boot is fine after a rolling deploy and terrible after an outage: a nudge about a message from three days ago is noise, not help, and a dozen of them arriving together is how a team decides the bot is a nuisance. The drop takes the same claim a delivery takes, so the row goes exactly once fleet-wide — and a nudge is never both dropped here and sent elsewhere.
See the private-nudge rule for why this is the single most important design decision in the product.
Every message action, and what it maps to
ChannelAction | Visibility | Used for |
|---|---|---|
Ephemeral | author only | a nudge with no rewrite suggestion |
EphemeralWithButtons | author only | a nudge with Post this / Dismiss |
ThreadReply | public, in-thread | context cards, accepted rewrites |
Post | public, in-channel | the charter intro when the bot is invited |
DirectMessage | private DM | the digest |
Replace | replaces the interactive message | button acknowledgements |
Each one increments ever_async_actions_total{channel,kind}.
What the pipeline deliberately does not do
- It never posts on the author's behalf. The rewrite is a suggestion; the author presses the button. Auto-posting was cut from the MVP on purpose.
- It never nudges twice. Four guards on the same failure: dedup at ingress, a
single scheduler slot per message key within a process, an atomic
claim_pending_nudgeacross processes, and cancel-on-follow-up. - It never classifies bot traffic. Including its own.
- It never blocks a webhook. Slow work is always on a spawned task.
Where this lives in the code
| Stage | Source |
|---|---|
| Ingress + routing | crates/server/src/ingress.rs |
| Everything else | crates/core/src/pipeline.rs |
| Heuristics | crates/core/src/classify.rs |
| Reference extraction | crates/core/src/refs.rs |
| Nudge timer (per process) | crates/core/src/scheduler.rs |
| Nudge delivery claim (fleet-wide) | claim_pending_nudge in crates/core/src/storage.rs + each backend |
| Restore on boot + staleness cutoff | Pipeline::start_background in crates/core/src/pipeline.rs, called from run_with_sso in crates/server/src/lib.rs |
| Rendering | crates/core/src/action.rs |