Nexus explainer · companion to Stack by Example
Hermes is the small messaging service every Nexus component uses to raise an alert. A service makes one HTTP call. Hermes stores the alert and answers 202 Accepted, and a separate loop posts it to Slack later.
The slow and unreliable step, calling Slack, never happens while the caller waits. It runs after the 202, in a different container, driven from the database. The database commit is the hand-off.
nexus-trader/…/notify.py::alert · nexus-data-api/common/hermes_notify.py · system-monitor/controller/alerts.py · nexus-backup/scripts/unit_alert.py
Nine callers across the estate raise alerts. Each sends a message with a tenant, client, idempotency key, a Slack route and text.
hermes/src/hermes_api/main.py::submit_message
POST /v1/messages checks a bearer key scoped to one tenant and app, stores the message, and returns 202 with a status URL. A duplicate gets 202 with the original message id.
hermes/src/hermes_api/store.py::SqlAlchemyStore.submit_message
SQLAlchemy 2.0 over Postgres. One commit writes the message, one delivery per recipient and channel, their outbox rows and an event. A unique constraint on tenant, client and key makes submits idempotent.
hermes/src/hermes_api/maintenance.py::run_maintenance
A second container runs the same image in a 60-second loop: send the daily digest, re-read pending deliveries, post them, then run webhook callbacks. Alert volume is low and a minute of latency is fine, so a loop over a Postgres table needs no broker to run or monitor. RabbitMQ support was scaffolded, but production uses the database path.
hermes/src/hermes_api/worker.py::SlackRouteDispatcher.send, DeliveryWorker.run_once
One incoming webhook per route (ops, watchlist, digest, planner). The route comes from the payload, then the recipient, then defaults to ops.
hermes/docker-compose.yml · .github/workflows/deploy.yml
Docker Compose on its own VPS: API, maintenance, Postgres 16. CI tests, builds and pushes to GHCR, then pulls and restarts over SSH. Callers reach it over the private network.
Every caller sends inline and synchronously, with a timeout, and none lets an alert failure become its own failure. They differ in what happens when Hermes is unreachable.
| Caller | Sends with | Timeout | If Hermes fails | Idempotency key |
|---|---|---|---|---|
| Tradernotify.py::alert | stdlib urlopen, in the calling thread | 10 s | silent returns False, never raises | app:dedupe_key, or a random key per event |
| Data APIhermes_notify.py | requests.post | 5 s | fallback logs a warning, then posts straight to Slack | random per event |
| System monitorcontroller/alerts.py | urlopen, inside the ingest request | 10 s each | logged error in the log; the ingest continues | host-app-outcome-day |
| Unit failure handlernexus-unit-alert@.service | a separate oneshot process started by OnFailure= | 10 s | loud exits 1, so the handler itself shows as failed | unit-failure:unit:hour |
| Arbitrage enginecoordinator.py::_alert_execution | blocking urlopen inside an async def | 15 s | blocks the event loop while it waits | per execution |
The key decides whether a repeated alert becomes one Slack post or many. Hermes keeps the first message for a key and answers every repeat with the same id.
unit-failure:nexus-foo.service:2026-09-30T14A unit failing every 15 minutes pages once an hour. A failure that persists re-alerts in the next hour.
nexus-trader:3f9c1a7e0b2d4c55Every event is a new message. Right for one-off events, noisy for anything that repeats.
nexus-trader:auto-halt:{account}Only the first ever is delivered, because keys never expire. Adding a time bucket fixes it.
Delivery is at least once. Exactly-once to Slack isn't possible unless Slack de-duplicates, so the design aims for rare duplicates and stable keys upstream.
| Situation | What the caller sees | What reaches Slack |
|---|---|---|
| Hermes is down | Waits up to its timeout, then carries on. The trader stays silent, the data API falls back to Slack, and the unit handler fails loudly. | Only the data API's direct fallback. |
| Slack is down | Nothing. It already has its 202. | Nothing. The delivery is marked failed and is not retried. |
| Same alert 100 times | 100 × 202, all with one message id when the key is stable. | One post with a stable key, 100 with a random key. |
| Loop crashes after posting, before marking | Nothing. | The row is still pending, so the next sweep posts it again. That is at least one attempt, with a possible duplicate. |
| Either container restarts | A timeout if the API was down at that moment. | Nothing committed is lost; pending rows are re-read from Postgres. |
Tracing the code turned up these gaps. Each has a small, specific fix.
Callers still wait on the networkup to 10 s in a trading thread
FixPut alerts on a bounded in-process queue drained by a background thread, or write them to the caller's own outbox table. The trader already does claim, send and mark for one alert type.
Failed Slack posts are never retriedfailed is terminal; the dead-letter status exists but is unused
FixRetry with exponential backoff up to a limit, then dead-letter and surface the count on the monitor.
Blocking call inside async codenexus-arb coordinator
FixUse an async HTTP client, or hand the call to a thread with run_in_executor.
Keys never expireconstant keys dedupe forever
FixTime-bucket every stable key, and add retention for old keys.
Health check doesn't check anything/health returns ok without touching the database
FixCheck the database connection and the age of the oldest pending delivery.
Two small bugsdaily digest posts twice; the API's in-process queue is never drained
FixQueue the digest from one source only, and remove the unused in-process publish from the API path.
Sources: hermes, nexus-trader, nexus-data-api, system-monitor, nexus-backup and nexus-arb main branches on 2026-09-29, traced file by file. Failure behaviour was confirmed in an offline simulation with the Slack call stubbed. No live endpoint was called.