Self-hosting
One static binary, one TOML file, embedded SQLite by default. everasync serve is a complete deployment — no database to provision, no queue, no external
service required. If you would rather it used Postgres (or MySQL), that is one
config line: see Choosing a storage backend.
[auth] mode defaults to local_trusted: no credential is required and every
caller is treated as a board principal. That is right for a private network and
wrong for anything reachable from the internet, because /api/v1 is writable.
Read Security & authentication before you publish a hostname.
Production hostnames
Ever Async runs on four hosts. The dashboard and the API are the same Rust binary — the dashboard is static files served by the same process.
| Role | Today (hyphen form) | Intended end state (dotted form) |
|---|---|---|
| Marketing website | async.ever.co | everasync.com |
| Dashboard (app) | app-async.ever.co | app.everasync.com |
| API + webhook ingress | api-async.ever.co | api.everasync.com |
| Documentation | docs-async.ever.co | docs.everasync.com |
Cloudflare Universal SSL and the zone's ACM packs cover ever.co and
*.ever.co only. A third-level host such as api.async.ever.co has no
certificate today and fails TLS. The dotted forms are the intended end state
and are already valid under everasync.com, so both are listed wherever hosts
appear.
Keep the base domain in one place per project — a constant or an
environment variable — so the switch is a single edit. This documentation site
does exactly that: one BASE_DOMAIN constant in docusaurus.config.ts.
Slack webhooks go to the API host:
https://api-async.ever.co/ingress/slack/events
https://api-async.ever.co/ingress/slack/commands
https://api-async.ever.co/ingress/slack/interactions
Path 1 — build from source
cargo build --release # rustup installs the pinned toolchain
./target/release/everasync init # writes everasync.toml into the CURRENT directory
$EDITOR everasync.toml # enable channels/connectors, set the *_env keys
cp .env.example .env # fill in secrets, export them into the env
./target/release/everasync serve # reads ./everasync.toml, listens on 0.0.0.0:8100
Trimming the binary
Every plugin ships in the default build; a section absent from the config leaves its plugin unregistered, so compiling one in costs nothing at runtime. For a purpose-built image, trim the feature list:
# Slack + GitHub + OpenRouter only
cargo build --release -p ever-async-cli \
--no-default-features --features slack,github,openrouter
Available features, by family:
| Family | Features |
|---|---|
| Channels | slack, discord |
| Connectors | github, jira, gauzy |
| LLM | anthropic, openai, openrouter, gemini, grok, ollama, compat |
| Notify | novu |
| Auth | auth-token, auth-local, auth-oauth |
| Storage | postgres, orm (and orm-mysql, which adds the MySQL driver) |
Every one of them is in the default set. memory and sqlite storage are not
features — they live in the core and are always available.
A config that names a provider whose feature was compiled out fails at startup
with a message listing what the binary does support — never silently. Two of
the auth features additionally back an everasync auth subcommand: auth token
needs auth-token, auth hash needs auth-local, and each says so by name if
you trimmed it out. auth-oauth backs no subcommand — the credentials it
consumes are issued by Slack, Discord, GitHub, Google or Ever Gauzy, not by this
binary. See Security & authentication.
Choosing a storage backend
[storage] driver picks one of four, and it is the decision that governs how
many replicas you can run:
driver | What it is | Replicas |
|---|---|---|
"sqlite" (default) | One file at [storage] path. Nothing to provision | One |
"postgres" | Several replicas share one database. url_env is required | Many (but read Scaling later) |
"orm" | The same storage over SeaORM: SQLite, Postgres or MySQL from one implementation, with versioned migrations. url_env picks the database by URL scheme; with no url_env it opens path as SQLite | Depends on the URL |
"memory" | Ephemeral — everything is lost on restart | Trying it out, tests |
# The default: nothing to run.
[storage]
driver = "sqlite"
path = "/app/data/everasync.db"
# Postgres. The connection string carries a password, so the file only names
# the variable holding it.
[storage]
driver = "postgres"
url_env = "EVERASYNC_DATABASE_URL" # postgresql://everasync:…@postgres:5432/everasync
pool_size = 16
An unknown driver refuses to start and lists the ids this binary can
actually open, rather than quietly falling back to a local file. Switching an
existing install to driver = "orm" is one word — the same everasync.db is
adopted and migrated in place, and the switch is reversible. Full key reference:
Configuration → [storage].
The dashboard
Optional when building from source. To serve it, build it once and keep
packages/apps/dashboard/dist relative to the server's working directory —
that is exactly where the static fallback looks:
pnpm install && pnpm --filter @ever-async/dashboard build
Without it, unrouted paths answer 404 with the build command in the body
rather than a bare 404 — the API, /healthz and /ingress/* are unaffected.
The status stays 404 on purpose: that same fallback catches a mistyped
/api/v1/..., and a 200 there would report success for a path that does not
exist.
os error 10013 on bindIf serve fails to bind, the port sits inside a Hyper-V/WinNAT excluded port
range — commonly 8080–8579, which includes the default 8100. Windows
reserves those ranges for NAT and refuses the bind with a permissions error
rather than "address in use", which is why it looks like a firewall problem.
netsh interface ipv4 show excludedportrange protocol=tcp
Pick a [server] bind port outside every listed range.
Path 2 — Docker
The repo Dockerfile builds the binary and the dashboard into one image
(binary plus packages/apps/dashboard/dist under /app, non-root user, port
8100):
docker build -t everasync .
docker run -d --name everasync \
-p 8100:8100 \
--env-file .env \
-v "$PWD/everasync.toml:/app/everasync.toml:ro" \
-v "$PWD/data:/app/data" \
everasync
Point [storage] path at the mounted volume so the database survives container
replacement:
[storage]
path = "/app/data/everasync.db"
The config is mounted read-only and holds no secrets — those arrive through
--env-file. That separation is the whole reason for the
*_env rule.
Path 3 — docker compose
cp .env.example .env # fill in secrets
everasync init # or write everasync.toml by hand
docker compose up -d
docker-compose.yml is a single service: build: ., 8100:8100,
env_file: .env, ./data mounted at /app/data for the SQLite file, and
everasync.toml mounted read-only.
Path 4 — Kubernetes
Nothing special is required — one Deployment, one Service, one Ingress. Four things to get right:
1. State. With the default driver = "sqlite" the database is a file: use a
PersistentVolumeClaim and one replica.
spec:
replicas: 1 # see "Scaling later" below
strategy:
type: Recreate # never two pods on one RWO volume
With driver = "postgres" there is no volume to mount and no Recreate strategy
needed. Nudges are safe above one replica — delivery is claimed in the database,
not held in a process, and every pod restores the whole pending queue on start —
but sessions (local/oauth auth) and the digest loop are still
per-process, so read Scaling later before raising replicas.
2. Config and secrets, separately. The TOML in a ConfigMap, the values in
a Secret — the *_env rule maps onto Kubernetes exactly:
volumeMounts:
- name: config
mountPath: /app/everasync.toml
subPath: everasync.toml
readOnly: true
envFrom:
- secretRef:
name: everasync-secrets # SLACK_BOT_TOKEN, GITHUB_TOKEN, EVERASYNC_API_TOKENS, …
3. Probes. /healthz returns ok and touches nothing expensive:
livenessProbe:
httpGet: { path: /healthz, port: 8100 }
readinessProbe:
httpGet: { path: /healthz, port: 8100 }
Never point a probe at /api/v1/... or / — a probe that does real work turns
a slow dependency into a restart loop.
Scrape /metrics with a ServiceMonitor or a
prometheus.io/scrape annotation.
4. Authentication, if the Ingress is reachable from outside the cluster.
An Ingress object is usually a public hostname, and the API is writable, so
this is the deployment path where the default matters most:
[auth]
mode = "authenticated"
provider = "token"
[auth.providers.token]
tokens_env = "EVERASYNC_API_TOKENS"
EVERASYNC_API_TOKENS goes in the same Secret as everything else — the
*_env rule is unchanged. Mint the values with everasync auth token. The
token provider is stateless, unlike the local and oauth providers'
in-memory sessions, so it is also the one that survives a pod restart cleanly.
For humans rather than machines, provider = "local" gives the dashboard a
password form, and provider = "oauth" gives it browser SSO against Slack,
Discord, GitHub, Google and Ever Gauzy — nothing to mint and no passwords to
store. Follow Setting up SSO for the exact redirect URIs
and scopes.
subPath ConfigMap mount never hot-reloadsEver Async reads everasync.toml once, at startup, and the standard mount above
is a subPath mount — which the kubelet resolves once when the container starts
and never refreshes. Editing the ConfigMap changes nothing in the running
pod. kubectl rollout restart deployment/everasync after every config change.
Leave /ingress/* and /healthz out of any nginx.ingress.kubernetes.io/auth-*
annotation you add at the Ingress: Slack cannot satisfy a basic-auth challenge
and a probe that gets a 401 restarts the pod. Full treatment in
Security & authentication.
Reverse proxy / TLS
Slack only delivers webhooks over HTTPS, and Ever Async serves plain HTTP
on :8100. Terminate TLS in front of it — Caddy (reverse_proxy localhost:8100,
automatic certificates), nginx + certbot, Traefik, or a Cloudflare Tunnel
pointed at http://localhost:8100.
Signature verification is computed over the raw request body. A proxy that
re-encodes, pretty-prints, re-chunks or otherwise rewrites the body invalidates
every signature and every webhook fails with Verification errors — which look
exactly like a wrong signing secret. This is the single most common deployment
failure.
Also make sure the host clock is right: the timestamp check that guards against replay attacks rejects requests that look stale, and clock skew reads as skew.
Two more:
- Only
/ingress/*needs to be reachable from the internet. Slack must reach it; nothing else must. (Ever Async's own auth allowlist is six prefixes —/ingress,/healthz,/api/v1/auth/mode,/api/v1/auth/loginand the two SSO browser routes/api/v1/auth/oauth/*/{start,callback}— but that is the app's gate, not a reason to publish them. If you use SSO, the callback route does have to be reachable by the person's browser, which is what[auth.providers.oauth] base_urlnames.) /api/v1is a writable API, and authentication is off by default.PUT /api/v1/charterrewrites the rules every nudge cites,PUT /api/v1/policies/…can silence the bot in a channel, andPOST /api/v1/digest/sendsends real messages to real people./metricsis not public either. Ingress is different — it is protected by each platform's request signature, which is why it stays anonymous.
On a private network the default ([auth] mode = "local_trusted", no
credential required) is correct and deliberate. Reachable from the internet it
means anyone who knows the hostname can turn your bot off.
Turn on the optional [auth] block — one
config change with the token provider — and read
Security & authentication for the threat model, the
role-to-endpoint table, and a lockdown checklist if you are exposed right now.
Backups
Everything that matters is in the storage backend: events, policies, the stored charter, pending nudges, context records, and — once anyone signs in — tenants, memberships and workspace bindings.
On sqlite (and orm pointed at a file) that is the single file at
[storage] path. Copy it on a schedule, and use the SQLite backup API rather
than cp on a live file:
sqlite3 /app/data/everasync.db ".backup '/backup/everasync-$(date +%F).db'"
On postgres (and orm with a server URL) Ever Async is an ordinary
database client: your existing pg_dump / WAL-archiving story covers it, and
there is nothing on the container's filesystem worth keeping.
Losing it is not catastrophic — Ever Async re-learns from traffic — but you lose per-conversation policies, the charter, and every tenant membership, so a restore is cheaper than a re-setup.
Scaling later
SQLite is the default, not the only option, and it is a deliberate one:
"easy to run" beats scale most self-hosters do not have. The Storage trait is
the exit, and it has been used — memory, sqlite, postgres and orm all
implement the same interface, tenant column included, so moving off SQLite is one
config line and no code.
Before running more than one replica, four things must be true. Two are solved, two are not:
-
Storage is not a file. SQLite on a shared volume is not a multi-writer database. ✅ Solved — set
driver = "postgres"(or"orm"with a Postgres/MySQL URL). Idempotent event claiming is a database-level operation, so de-duplication stays correct across replicas. -
The nudge scheduler. ✅ Solved. It is still an in-process
tokiodelay map, and it always will be — timers cannot elect anything across processes. What changed is that they no longer have to. Before sending, the fire path claims the nudge in storage, and a claim is a single atomic statement (DELETE … WHERE tenant = ? AND message_key = ?, its affected-row count is the answer) that exactly one caller in the fleet wins. So N replicas arm N timers for the same nudge, all N fire, one delivers and the rest quietly do nothing — including the big fan-out at boot, where every replica restores every pending nudge.Two consequences worth knowing:
- The claim consumes the row. A cancelled nudge — a Dismiss button, an
/async off, an author follow-up — leaves no row, so a timer surviving on another replica finds nothing to win and stays silent. That was the second double-nudge path, and the nastier one: neither a button press nor/async offleaves an author follow-up behind, so the old self-correction re-check said "still warranted" and nudged someone who had just asked it to stop. - Delivery is at-most-once. The claim is taken immediately before the send, so a nudge whose delivery fails is gone rather than retried. That is deliberate: putting the row back would turn a delivery that timed out but actually landed into a second nudge hours later. Missing one nudge is a bad minute; nudging twice is why someone turns the product off.
Watch
ever_async_nudges_claim_lost_total. It should be ~0 at one replica and grow withreplicas - 1per nudge above that — it is the shape of the fan-out, not an error. - The claim consumes the row. A cancelled nudge — a Dismiss button, an
-
Sessions. ⚠️ Still open for
localandoauth. Both keep their session store in process memory, so a second replica does not see the first one's sessions and a restart logs everyone out. Thetokenprovider is stateless and unaffected. -
The digest loop. ⚠️ Still open.
spawn_digest_loopis a plain per-processtokio::time::interval, so N replicas deliver N digests per interval.everasync servestarts it — but only when[digest] enabled = true, and the default isfalse, so a scaled-out deployment that never enabled it is unaffected. Keep it disabled whilereplicas > 1, or trigger digests from a single external cron againstPOST /api/v1/digest/send.
So today, with driver = "postgres" (or "orm" on a Postgres URL): more than
one replica is safe for messages and nudges, and blocked on sessions if you
use the local or oauth auth providers, and on the digest if you have enabled
it. With driver = "sqlite" it is still one replica, because the file is.
Restarts and the pending queue
A nudge sits out its grace period on an in-process timer, so a restart in the
middle of that window used to lose it: the row was persisted, and nothing read it
back. everasync serve now re-arms the whole pending set at boot, and the
delivery claim keeps that safe at any replica count — every replica restores
every nudge, exactly one delivers.
What boot has to judge is the nudges whose grace period expired while nothing was
running. [nudges] max_overdue_secs (default 900, 15 minutes) is the cutoff:
anything more overdue than that is dropped rather than delivered, because a nudge
about a message from hours ago is noise, and a batch of them arriving together
the moment a pod comes back is worse than the restart was.
[nudges]
max_overdue_secs = 900 # 0 = deliver nothing that fell due while stopped
Dropped nudges are consumed through the same claim a delivery uses — the row goes once fleet-wide, not once per replica — and counted:
increase(ever_async_nudges_stale_dropped_total[1d])
Zero after a rolling deploy. Non-zero after an incident, and then it is the count of messages that never got their nudge because of it.