nexus-data-api.service
Type=simple ExecStart=uvicorn … --workers 1 Restart=on-failure RestartSec=10
The data API runs as its own service user. The research API and the other FastAPI services follow the same pattern.
Nexus explainer · companion to Stack by Example
Nexus has no Celery, no Airflow and no application crontab. Every service and scheduled job on the production host is a systemd unit, and Postgres tables do the queueing a task broker would do. This page shows how that works and when I would choose something else.
Read from the production host on 2026-09-30 with systemctl, read-only.
The 15-minute intraday price fetch, as it runs today. The timer starts a oneshot service inside a cgroup. Output goes to journald, and a failure starts a second unit that raises an alert.
Type=simple ExecStart=uvicorn … --workers 1 Restart=on-failure RestartSec=10
The data API runs as its own service user. The research API and the other FastAPI services follow the same pattern.
Restart=always Wants=network-online.target After=network-online.target
The IG price capture. Identity variables are set after EnvironmentFile= so a .env can't override them.
ExecStart=… scheduler_tick.py --loop --interval-seconds 60 TimeoutStopSec=120
Replaced a 15-minute timer so the book cache stays warm between ticks. The old timer is disabled on deploy.
OnCalendar=Mon..Fri 17:30 Europe/London Persistent=false
Time-zone-aware calendar. Catch-up is off because a snapshot taken late would be wrong, not just late.
PathExistsGlob=…/queue/*.req
dev drops a request file and a root oneshot runs the job from an allow-list. A glob on *.req stops a junk file re-triggering the unit into its start limit.
ExecStart=… unit_alert.py %i TimeoutStartSec=60
One definition serves all 114 first-party services: each names it with OnFailure=nexus-unit-alert@%n.service.
ListenStream=127.0.0.1:… → systemd-socket-proxyd
Starts a proxy on first connection, because the private-network address may not exist yet at boot.
Restart=always Slice=ci.slice (MemoryMin=8G)
18 self-hosted CI runners under the dev user, with a guaranteed memory floor.
systemd-run --user --scope --slice=agentwork.slice -p MemorySwapMax=0 -p MemoryMax=…
Heavy one-off jobs get their own cgroup with a memory ceiling, created on the fly.
No overlapping runsbuilt in
A timer that fires while its service is still running is absorbed. On 30 September an hourly price job had been running 23 minutes and its timer simply showed no next run. Dispatchers that loop add flock -n.
Catch-up after downtimePersistent=true
A missed run fires once at boot. On 22 August, 25 catch-up runs fired before Postgres was up and failed, so 64 units now carry an After=postgresql drop-in.
Failure pagingOnFailure=
Every first-party service names the alert template, enforced by a test in 9 repositories. The handler has no OnFailure= of its own, so a messaging outage can't loop.
Expected non-successSuccessExitStatus=75
When a vendor throttles, the fetch exits 75 by design. systemd counts it as success, so it doesn't page anyone.
Deploy awarenessExecStartPre · ExecCondition
37 data jobs wait while a deploy is running, and 2 skip entirely.
Spreading loadRandomizedDelaySec · OnCalendar offsets
13 timers add random delay. The capture backup runs at :05, :20, :35 and :50, never on the quarter-hours that trading and the IG rate limiter use.
Resource isolationcgroup v2
CPUWeight=20 keeps backtests behind trading. Postgres has an 18 GB memory floor and no ceiling, because the database must not be the thing that dies. Agent jobs run with swap off.
Least privilegehardening directives
Trader units use ProtectSystem=strict with one writable path, plus NoNewPrivileges and namespace restrictions. The scheduler omits PrivateTmp on purpose, because the operator's kill-switch file lives in a shared temporary directory.
Logsjournald
Every unit logs to the journal under its own identifier, kept 30 days within a 16 GB cap.
Celery needs a broker because workers must claim tasks without colliding. Postgres already does that with row locks, so Nexus uses the database it already runs.
Backtest queuenexus-sim-api/src/nexus_sim/db.py
UPDATE … WHERE run_id = (SELECT … FOR UPDATE SKIP LOCKED LIMIT 1) RETURNING. Workers poll every 20 s when idle. A supervisor keeps N spawned worker processes alive inside one Restart=always service, and resets orphaned runs at start, which is safe because systemd guarantees one instance.
API jobsnexus-data-api/app/routers/jobs.py
Eight endpoints return 202, write a job_state row and start a thread. Clients poll GET /jobs/{id}.
In-process schedulernexus-janus/app/scheduler.py
APScheduler with two cron triggers inside one small service. It suits a job that only matters while that service is up.
| cron | systemd timers | Celery | Airflow | |
|---|---|---|---|---|
| What it is | Runs a command at set times | Timer units that start service units under the init system | Distributed task queue: producers, a broker and workers | Workflow orchestrator for DAGs of tasks |
| Triggers | Five-field time spec | Calendar with seconds and time zones, intervals after boot or last run, file paths, sockets | Code enqueues tasks; beat adds periodic ones | A schedule or a dataset event per DAG |
| Overlap | Starts again regardless; you add flock | A fire while running is absorbed | Needs a lock | max_active_runs and concurrency settings |
| Missed runs | Lost (anacron covers daily jobs) | Persistent=true runs once at boot | No backfill | Catch-up and backfill of every missed interval |
| Failures | Output mailed to MAILTO | Exit codes, Restart=, OnFailure= to any unit | Per-task retries with backoff | Retries, SLAs and callbacks per task |
| Dependencies | None | Start ordering (After, Requires), not data flow | Chains, groups and chords | A DAG of tasks, its main strength |
| Resource limits | None per job | Per-unit cgroup: memory, CPU weight, task count | Worker concurrency and prefetch | Pools and queues; limits depend on the executor |
| Visibility | Syslog and mail | journald per unit, list-timers, systemctl status | Flower, as an add-on | Web UI with run history and task logs |
| Scale-out | One host | One host | Many workers across hosts | Many workers through its executors |
| Extra infrastructure | None | None: systemd is already PID 1 | A broker, workers and a result backend | Scheduler, web server, metadata database, workers |
| In Nexus | Not used. Cron jobs would run outside the agents' memory cap, so the rule is user timers | 62 timers, 99 services, 2 path units | Replaced by Postgres SKIP LOCKED queues | Not needed at this size |
One host, tens of independent jobs, Postgres already running, and a need for cgroup control next to latency-sensitive trading. That describes Nexus today.
When jobs form real data dependencies across many steps, and people need run history, date-based backfills and a UI. Vendor ingestion spread across several hosts would qualify.
When many short tasks must fan out across machines, or request handling needs a worker pool beyond one host. Celery fits, as do lighter options such as RQ or Dramatiq.
nproc --all.systemd-analyze security scores it 9.2 (unsafe), against 5.8 for the price capture.Nice=10 from the CI runner units, since it does nothing under cgroup v2.Sources: systemctl cat, show and list-timers on the production host (read-only, 2026-09-30); unit files on each repository's main branch; handbook platform notes, ADR-0008 and ADR-0009, and the unit-alert README.