London Data

Nexus explainer · companion to Stack by Example

Scheduled by systemd

Nexus has no Celery, no Airflow and no application crontab. Every service and scheduled job on the production host is a systemd unit, and Postgres tables do the queueing a task broker would do. This page shows how that works and when I would choose something else.

Services9926 long-running, 73 oneshot jobs
Timers62active, most with Persistent=true
Failure alerts114/114first-party services with OnFailure=
Path units2deploy queue and job runner
CI runners18user services in their own slice
Celery · Airflow0APScheduler in one small service

Read from the production host on 2026-09-30 with systemctl, read-only.

Anatomy of one scheduled job

The 15-minute intraday price fetch, as it runs today. The timer starts a oneshot service inside a cgroup. Output goes to journald, and a failure starts a second unit that raises an alert.

Timer …-intraday-5min.timer OnCalendar=*:0/15:20 Persistent=true Fires while the job still runs? The fire is absorbed. No overlap. CGROUP · CPUWeight, memory limits Service (Type=oneshot) After=postgresql (drop-in) ExecStartPre: wait if deploying SuccessExitStatus=75 journald SyslogIdentifier= 30 days, 16 GB cap nexus-unit-alert@%n oneshot · unit_alert.py no OnFailure of its own Hermes 202 Accepted key: unit + hour Slack ops route starts stdout OnFailure= exit 75 is a stand-down: no page POST later
Each concern a job framework would give you is a directive here: schedule and catch-up on the timer; ordering, deploy-awareness and exit-code meaning on the service; resource limits on the cgroup; logs in journald; paging through a separate template unit.

The unit kinds in use

Long-running API

nexus-data-api.service

Type=simple
ExecStart=uvicorn … --workers 1
Restart=on-failure
RestartSec=10

The data API runs as its own service user. The research API and the other FastAPI services follow the same pattern.

Long-running daemon

nexus-trader-capture.service

Restart=always
Wants=network-online.target
After=network-online.target

The IG price capture. Identity variables are set after EnvironmentFile= so a .env can't override them.

Loop service

nexus-trader-scheduler.service

ExecStart=… scheduler_tick.py
  --loop --interval-seconds 60
TimeoutStopSec=120

Replaced a 15-minute timer so the book cache stays warm between ticks. The old timer is disabled on deploy.

Timer + oneshot

nexus-trader-target-snapshot.timer

OnCalendar=Mon..Fri 17:30 Europe/London
Persistent=false

Time-zone-aware calendar. Catch-up is off because a snapshot taken late would be wrong, not just late.

Path unit

nexus-jobrunner.path

PathExistsGlob=…/queue/*.req

dev drops a request file and a root oneshot runs the job from an allow-list. A glob on *.req stops a junk file re-triggering the unit into its start limit.

Template

nexus-unit-alert@.service

ExecStart=… unit_alert.py %i
TimeoutStartSec=60

One definition serves all 114 first-party services: each names it with OnFailure=nexus-unit-alert@%n.service.

Socket activation

nexus-risk-app-loopback.socket

ListenStream=127.0.0.1:…
→ systemd-socket-proxyd

Starts a proxy on first connection, because the private-network address may not exist yet at boot.

User services

github-runner-*.service

Restart=always
Slice=ci.slice   (MemoryMin=8G)

18 self-hosted CI runners under the dev user, with a guaranteed memory floor.

Transient scope

run-capped

systemd-run --user --scope
  --slice=agentwork.slice
  -p MemorySwapMax=0 -p MemoryMax=…

Heavy one-off jobs get their own cgroup with a memory ceiling, created on the fly.

Framework features, done with directives

Where Postgres replaces the broker

Celery needs a broker because workers must claim tasks without colliding. Postgres already does that with row locks, so Nexus uses the database it already runs.

cron, systemd, Celery and Airflow

cronsystemd timersCeleryAirflow
What it isRuns a command at set timesTimer units that start service units under the init systemDistributed task queue: producers, a broker and workersWorkflow orchestrator for DAGs of tasks
TriggersFive-field time specCalendar with seconds and time zones, intervals after boot or last run, file paths, socketsCode enqueues tasks; beat adds periodic onesA schedule or a dataset event per DAG
OverlapStarts again regardless; you add flockA fire while running is absorbedNeeds a lockmax_active_runs and concurrency settings
Missed runsLost (anacron covers daily jobs)Persistent=true runs once at bootNo backfillCatch-up and backfill of every missed interval
FailuresOutput mailed to MAILTOExit codes, Restart=, OnFailure= to any unitPer-task retries with backoffRetries, SLAs and callbacks per task
DependenciesNoneStart ordering (After, Requires), not data flowChains, groups and chordsA DAG of tasks, its main strength
Resource limitsNone per jobPer-unit cgroup: memory, CPU weight, task countWorker concurrency and prefetchPools and queues; limits depend on the executor
VisibilitySyslog and mailjournald per unit, list-timers, systemctl statusFlower, as an add-onWeb UI with run history and task logs
Scale-outOne hostOne hostMany workers across hostsMany workers through its executors
Extra infrastructureNoneNone: systemd is already PID 1A broker, workers and a result backendScheduler, web server, metadata database, workers
In NexusNot used. Cron jobs would run outside the agents' memory cap, so the rule is user timers62 timers, 99 services, 2 path unitsReplaced by Postgres SKIP LOCKED queuesNot needed at this size

When I would choose something else

Stay on systemd

One host, tens of independent jobs, Postgres already running, and a need for cgroup control next to latency-sensitive trading. That describes Nexus today.

Move to Airflow

When jobs form real data dependencies across many steps, and people need run history, date-based backfills and a UI. Vendor ingestion spread across several hosts would qualify.

Move to a task queue

When many short tasks must fan out across machines, or request handling needs a worker pool beyond one host. Celery fits, as do lighter options such as RQ or Dramatiq.

Lessons written down along the way

What I would change

Sources: systemctl cat, show and list-timers on the production host (read-only, 2026-09-30); unit files on each repository's main branch; handbook platform notes, ADR-0008 and ADR-0009, and the unit-alert README.