Sensitivity and the precision-over-recall stance
Sensitivity is one setting with three values, and it answers exactly one question: how much of the ambiguous middle does Ever Async act on?
[policy]
sensitivity = "conservative" # default · "balanced" · "high"
The three levels
| Level | Acts on | LLM consulted | Confidence bar |
|---|---|---|---|
conservative | textbook heuristic hits only — naked greeting, artifact-without-link | no | — |
balanced | the above, plus AI verdicts | yes | ≥ 0.80 |
high | the above, plus weaker AI verdicts | yes | ≥ 0.60 |
conservative is the default, and it is the only level that never spends a
token: the deterministic rules are the entire decision. Under balanced and
high, only messages the heuristics were not certain about ever reach a
model — a clear pass and a textbook fail are both decided locally, in
microseconds.
Change it per conversation at any time:
/async sensitivity balanced
Why the default is the timid one
The product's riskiest assumption is that people tolerate software commenting on their communication at all. Given that, the two kinds of mistake are not remotely symmetric:
| Mistake | What it costs |
|---|---|
| False negative — a genuinely vague message is missed | Nothing new. That message was going to be vague anyway; the world is exactly as it was before Ever Async was installed. |
| False positive — a perfectly good message is flagged | Trust. The author learns the tool does not understand their work, tells a colleague, and the channel turns it off. |
A false negative is a missed improvement. A false positive is an active withdrawal. You cannot average them.
So the stance is precision over recall, stated plainly in the code:
/// Only flag textbook cases (naked greeting, artifact with no link).
/// Default: false positives are the product's #1 kill risk.
#[default]
Conservative,
What "certain" means
The heuristic classifier returns a verdict and a certain flag. certain
is not "high confidence" — it means a textbook rule fired, or the message
clearly cleared the bar:
| Message | Verdict | certain | What happens |
|---|---|---|---|
Hi team! | LowContext (0.95) | ✅ | acted on at every level |
I finished the PR, can someone review it? | LowContext (0.90) | ✅ | acted on at every level |
| a message with a link, or ≥ 30 words | Clear (0.90) | ✅ | never acted on |
any update on this? | Clear (0.55) | ❌ | the ambiguous middle |
That last row is the only row sensitivity affects. Under conservative it is
left alone. Under balanced or high it goes to the AI provider, and the
provider's verdict is used only if it says low_context and clears the
level's confidence bar.
An AI failure is never a flag
If the provider errors, times out, rate-limits, or answers something that is not the demanded JSON shape, the error is logged and the heuristic verdict stands. The pipeline reads an LLM error as "no opinion" — never as "probably bad". That is why the parsers refuse to guess: a lenient parser that salvaged half a malformed reply would turn a model outage into a wave of false positives.
The same applies to the fallback chain. A provider that is down is skipped, the next one answers, and if every one of them fails the message is simply judged by heuristics. See AI providers.
Choosing a level
Start on conservative. Run for a week and read the counters
(Operations):
ever_async_nudges_total— how often it actsever_async_nudges_cancelled_total— how often the author self-corrected first (good: the grace window is doing its job)ever_async_rewrites_accepted_total— how often a suggestion was worth one click (the honest quality signal)
Move to balanced when people are asking for more — when the complaint is
"it missed one" rather than "it flagged mine". high is for a team that has
explicitly opted into an aggressive coach; it is not a good default for a
channel with newcomers in it.
Sensitivity is per conversation, so a single deployment can run high in the
team's own engineering channel and conservative everywhere else.
What sensitivity does not change
- Nudges stay private at every level. See the private-nudge rule.
- The grace window still applies.
highdoes not mean faster. - Context cards are unaffected. They are gated by
[policy] min_confidenceon the connector's confidence, which is a different number entirely — see policy.