London Data

Nexus explainer · companion to Stack by Example

The Hermes Alert Path

Hermes is the small messaging service every Nexus component uses to raise an alert. A service makes one HTTP call. Hermes stores the alert and answers 202 Accepted, and a separate loop posts it to Slack later.

Where the non-blocking happens

The slow and unreliable step, calling Slack, never happens while the caller waits. It runs after the 202, in a different container, driven from the database. The database commit is the hand-off.

Caller side: bounded and fail-silentA synchronous HTTP POST with a 5 to 15 second cap. It never raises. If Hermes is down, the caller loses at most its timeout and carries on.
Hermes side: asynchronousOne database commit, then 202. A maintenance loop picks up pending deliveries within about a minute and posts them to Slack.

One alert, end to end

ASYNCHRONOUS FROM HERE · NEXT SWEEP WITHIN ABOUT 60 s Trader notify.alert() Hermes API FastAPI, 1 worker Postgres same compose stack Maintenance loop same image, sleep 60 Slack incoming webhooks waits ≤ 10 s POST /v1/messages bearer key · idempotency key one transaction message · deliveries · outbox · event 202 Accepted message_id · status_url caller carries on read pending up to 200 deliveries POST webhook route · 10 s timeout 2xx or error mark delivery succeeded or failed
The caller's only wait is one round trip to Hermes and one database commit. Everything that touches Slack happens inside the shaded band, driven by rows in Postgres, so a restart of either container loses nothing that was committed.

The parts

Callers

nexus-trader/…/notify.py::alert · nexus-data-api/common/hermes_notify.py · system-monitor/controller/alerts.py · nexus-backup/scripts/unit_alert.py

Nine callers across the estate raise alerts. Each sends a message with a tenant, client, idempotency key, a Slack route and text.

Hermes API

hermes/src/hermes_api/main.py::submit_message

POST /v1/messages checks a bearer key scoped to one tenant and app, stores the message, and returns 202 with a status URL. A duplicate gets 202 with the original message id.

Store

hermes/src/hermes_api/store.py::SqlAlchemyStore.submit_message

SQLAlchemy 2.0 over Postgres. One commit writes the message, one delivery per recipient and channel, their outbox rows and an event. A unique constraint on tenant, client and key makes submits idempotent.

Maintenance loop

hermes/src/hermes_api/maintenance.py::run_maintenance

A second container runs the same image in a 60-second loop: send the daily digest, re-read pending deliveries, post them, then run webhook callbacks. Alert volume is low and a minute of latency is fine, so a loop over a Postgres table needs no broker to run or monitor. RabbitMQ support was scaffolded, but production uses the database path.

Slack routes

hermes/src/hermes_api/worker.py::SlackRouteDispatcher.send, DeliveryWorker.run_once

One incoming webhook per route (ops, watchlist, digest, planner). The route comes from the payload, then the recipient, then defaults to ops.

Deployment

hermes/docker-compose.yml · .github/workflows/deploy.yml

Docker Compose on its own VPS: API, maintenance, Postgres 16. CI tests, builds and pushes to GHCR, then pulls and restarts over SSH. Callers reach it over the private network.

How each caller sends

Every caller sends inline and synchronously, with a timeout, and none lets an alert failure become its own failure. They differ in what happens when Hermes is unreachable.

CallerSends withTimeoutIf Hermes failsIdempotency key
Tradernotify.py::alertstdlib urlopen, in the calling thread10 ssilent returns False, never raisesapp:dedupe_key, or a random key per event
Data APIhermes_notify.pyrequests.post5 sfallback logs a warning, then posts straight to Slackrandom per event
System monitorcontroller/alerts.pyurlopen, inside the ingest request10 s eachlogged error in the log; the ingest continueshost-app-outcome-day
Unit failure handlernexus-unit-alert@.servicea separate oneshot process started by OnFailure=10 sloud exits 1, so the handler itself shows as failedunit-failure:unit:hour
Arbitrage enginecoordinator.py::_alert_executionblocking urlopen inside an async def15 sblocks the event loop while it waitsper execution

Three kinds of idempotency key

The key decides whether a repeated alert becomes one Slack post or many. Hermes keeps the first message for a key and answers every repeat with the same id.

Time-bucketed preferred

unit-failure:nexus-foo.service:2026-09-30T14

A unit failing every 15 minutes pages once an hour. A failure that persists re-alerts in the next hour.

Random no dedupe

nexus-trader:3f9c1a7e0b2d4c55

Every event is a new message. Right for one-off events, noisy for anything that repeats.

Constant dedupes forever

nexus-trader:auto-halt:{account}

Only the first ever is delivered, because keys never expire. Adding a time bucket fixes it.

What happens when something breaks

Delivery is at least once. Exactly-once to Slack isn't possible unless Slack de-duplicates, so the design aims for rare duplicates and stable keys upstream.

SituationWhat the caller seesWhat reaches Slack
Hermes is downWaits up to its timeout, then carries on. The trader stays silent, the data API falls back to Slack, and the unit handler fails loudly.Only the data API's direct fallback.
Slack is downNothing. It already has its 202.Nothing. The delivery is marked failed and is not retried.
Same alert 100 times100 × 202, all with one message id when the key is stable.One post with a stable key, 100 with a random key.
Loop crashes after posting, before markingNothing.The row is still pending, so the next sweep posts it again. That is at least one attempt, with a possible duplicate.
Either container restartsA timeout if the API was down at that moment.Nothing committed is lost; pending rows are re-read from Postgres.

Known gaps and the next fix

Tracing the code turned up these gaps. Each has a small, specific fix.

Sources: hermes, nexus-trader, nexus-data-api, system-monitor, nexus-backup and nexus-arb main branches on 2026-09-29, traced file by file. Failure behaviour was confirmed in an offline simulation with the Slack call stubbed. No live endpoint was called.