London Data

Nexus explainer · companion to Scheduled by systemd

One Host, Many Tenants

One Linux server runs the market database, paper trading, ten APIs, backtests, 18 CI runners and a fleet of AI coding agents. The rule is simple: the database and trading keep running, whatever everything else does. This page shows how accounts, cgroups and scheduling enforce that, what happened when they didn't, and how the host is watched.

Host16 · 62 GiBCPU threads · RAM, plus 32 GiB swap
Service accounts10one per service, no login shell
Postgres floor18 GiBon Postgres and both parent slices; no ceiling
Agents + CI seat34 / 36 GiBMemoryHigh / MemoryMax, 12 of 16 CPUs
Swap for agent jobs0an overrun is killed at once, never swapped
Journal30 dayspersistent, capped at 16 GB

Read from the production host on 2026-09-30, read-only: cgroupfs, /proc, systemctl and journalctl.

Part 1 · Separating the workload

Accounts: who runs what

Each service runs as its own Linux user with no login shell, owns its own directory, and keeps its secrets in a file only it can read. The agents' account can read everything and change nothing directly.

AccountRunsCanCannot
market-dataData API, analysts, data timersWrite its own directory under /opt; its .env is mode 0600Log in; write system directories (ProtectSystem=full)
nexus-traderTrader API, price capture, schedulerWrite only /var/lib/nexus-trader (ProtectSystem=strict)Gain privileges (NoNewPrivileges)
nexus-sim-apiResearch API and backtestsRun spawned worker poolsTake CPU from trading (CPUWeight=20)
7 morepms, risk, arb, planner, RFQ simulator, monitor, research dataOne service eachLog in
postgresPostgreSQL 16 + TimescaleDBKeep an 18 GiB memory floor, held all the way up the treeHit a memory ceiling: it has none, by design
devAI agents, CI runners, toolingRead the system journal; read all database data (pg_read_all_data); queue deploys and jobssudo; change another service's files or secrets
adminAdministration by a personsudo (the only account with it)

How dev triggers a privileged action without sudo

devdeploy-now traderwrites a request file into a queue directory (ACL: dev rwx)
root · systemd.path unit firesDirectoryNotEmpty= or PathExistsGlob=*.req
root · dispatcherChecks before actingflock -n · name must be on a root-owned allow-list · refuses if that list is group- or world-writable
service userRuns the jobrunuser -u <user> · env -i · timeout --kill-after=30 · args checked by regex, ≤ 4096 bytes

Only a name crosses the boundary. Everything that runs is decided by root-owned configuration, so a compromised agent account can queue a deploy but can't choose what it executes.

cgroups: who gets protected, capped or weighted

cgroup v2 with the cpu, memory and pids controllers enabled. Services sit in system.slice; everything the agents and CI do sits under the dev user's slice, called the seat below, which carries the limits.

protect memory floor limit ceiling or quota weight share under contention
  • /controllers: cpu memory pids · io not enabled
    • system.sliceMemoryMin=18Gno ceiling, on purpose
      • nexus-*.service~24 long-running services, each as its own user
      • nexus-sim-backtest.serviceCPUWeight=20about 6% vs 25% of CPU under contention; every core when idle
      • pg-backup-offsite.serviceCPUWeight=10High 1G · Max 2Gswap 0
      • nexus-jobrunner-dispatch.serviceMax 16G · CPUQuota 80% · TasksMax 256
      • system-postgresql.sliceMemoryMin=18Gpasses the floor down
        • postgresql@16-mainMemoryMin=18G16G shared_buffers + 2G; no MemoryMax
    • user.slice
      • user-1001.slice(dev)Min 8GHigh 34G · Max 36GSwap 8GCPUQuota 1200%CPUWeight 20
        • user@1001.serviceMin 8Gdelegates cpu, memory, pids to the user manager
          • tmux-spawn-*.scopeagent sessions
          • ci.sliceMemoryMin=8G18 CI runners; 8G is their measured working set
          • agentwork.sliceMemorySwapMax=0heavy jobs, one capped scope each
          • app.slice / memguardMax 256M · swap 0the memory governor

Where the memory numbers come from

Postgres floor · 18 GiB Services 5 CI floor 8 Agents + CI seat · Max 36 GiB MemoryHigh 34 kernel 2 page cache floor 1 0 18 26 62 GiB
RAM to scale. The seat's ceiling is what remains after the floors: 61.9 − 18 − 5 − 2 − 1 ≈ 36 GiB. MemoryHigh sits 2 GiB below MemoryMax, so a job that drifts over gets throttled before it gets killed. Before 6 September the caps added up to more than the machine had.

Why 18 GiB for Postgres

The floor is the database's buffer pool plus a little headroom. It sounds generous until you see what it does and doesn't cost.

shared_buffers16 GiBabout 25% of RAM, the usual sizing
Headroom+2 GiBconnection processes
In use now19.2 GiB16.2 of it the buffer pool
Buffer hit ratio99.9%market_data, since stats began
Found and fixed while writing this page · 30 September 2026

A floor only counts if every parent has one

In cgroup v2 the kernel scales each level's protection by its parent's effective floor, and a parent at zero gives zero. Postgres has no memory ceiling of its own, so the only reclaim that reaches it is host-wide, the 6 September condition. Its two parent slices had no floor, so its 18 GiB was protecting nothing in exactly that case. The existing check read only Postgres's own value and stayed green.

cgroupmemory.min beforeafter
system.slice018 GiB
system-postgresql.slice018 GiB
postgresql@16-main18 GiB18 GiB
Effective under host-wide reclaim018 GiB

Root added a MemoryMin=18G drop-in to both slices at about 09:40 UTC, reloaded systemd and wrote the live values, without restarting the database. The remaining work is a check that reads every ancestor, with a negative control, so this can't pass silently again.

The other levers

In force

CPUWeight, not nice

CPUWeight=20   # backtests
CPUQuota=1200% # dev seat

Under cgroup v2, nice only orders processes inside one cgroup, so it can't protect one service from another. Weights share CPU only under contention, so backtests still use every core on a quiet box.

In force

Capped scopes for heavy jobs

run-capped --memory 8G -- <cmd>
→ systemd-run --user --scope
  --slice=agentwork.slice
  MemoryMax=8G MemoryHigh=80%

Each heavy job gets its own cgroup with swap off. A test allocation over the cap was killed with rc 137 and zero swap used. It warns when the request exceeds the seat's headroom.

Gate on · freezer observing

A memory gate on dispatch

LOADED if headroom < 10 GiB,
  MemAvailable < 10 GiB or
  PSI memory some avg60 ≥ 10%

One reader, memguard, decides whether new agent work may start. Dispatch refuses (exit 18) when closed. It can freeze running jobs with the cgroup freezer rather than kill them, but today it only logs. It covers the ground systemd-oomd would, which isn't installed; freezing rather than killing suits long agent jobs that can resume.

In force

Page-cache reclaim

echo "<bytes> swappiness=0" \
  > …/user@1001.service/memory.reclaim

A one-minute timer asks the kernel to drop the seat's page cache (up to 4 GiB per tick), so cache can't push the seat into throttling.

In force

/tmp is RAM

/tmp: 31G tmpfs
TMPDIR → on-disk path
tmp-reaper every 6 h

Files in a tmpfs count as shared memory. Dead session directories and orphaned test databases once filled 24 GiB of it, so work now writes to disk and a reaper clears leftovers.

In force

Serialisation and timeouts

flock -n /run/nexus-autodeploy.lock
timeout 2700s  # a deploy
lock_timeout=30s  # migration DDL

Root dispatchers skip if another pass holds the lock. Every job has a hard timeout. Migrations give up waiting for a lock rather than queueing everything behind them.

In force

Keeping clear of trading minutes

OnCalendar=*:5/15   # capture export
07:07, 19:07        # checker sweep

Trading and the broker rate limiter run on the quarter-hours, so other daily jobs avoid them. The rule came from a login collision on 31 August.

In force

Task limits

TasksMax=33%  # dev seat, 153,710
TasksMax=256  # job runner

A fork bomb or thread leak in the seat stops at its own limit. The seat has never reached it.

Not available

I/O isolation

io controller: not enabled
NVMe scheduler: none

IOWeight needs the io controller and ionice needs the BFQ scheduler; neither is in place. Disk contention is managed by memory floors that keep hot data cached instead.

When the host stopped responding

Each episode left a control behind. The timelines come from sar history (sysstat samples every 10 minutes), the persistent journal and cgroup event counters.

Swapped solid for 100 minutes

Three agent sessions ran mutation tests and noise experiments at once. Memory rose by about 21 GiB an hour. The seat had a MemoryMax but unlimited swap, and MemoryMax bounds RAM only, so the kernel swapped instead of killing anything. No OOM kill appears in that boot's log. Postgres had no floor, lost its cache and went to disk. ssh answered but couldn't fork a login, and the box needed a hardware reset.

Added: swap caps, capped scopes for heavy jobs, the 18 GiB Postgres floor, seat limits recomputed from the RAM budget, and a memory alert that was proven to deliver.

Page cache slowed CI

The seat sat at 33.2 of 34 GiB, mostly cache, and was throttled millions of times. A 40-minute CI run took over an hour.

Added: the page-cache reclaim timer.

/tmp full of memory

24 GiB of dead session directories, test output and 26 orphaned test Postgres servers sat in the RAM-backed /tmp. The seat cap held, so the database kept running, but ssh froze for about 15 minutes.

Added: an on-disk TMPDIR and the reaper.

Dispatches piled on

The seat's working set rose from 19 to 32 GiB in 15 minutes while new agent work kept arriving. ssh stalled and six jobs were frozen by hand.

Added: memguard, one memory verdict that dispatch, claims and drains all read.

The database, not the host

A deploy's migration waited for a lock behind a 57-minute query, and 45 connections queued behind the migration. The API was down for 33 minutes while the host was fine.

Added: lock_timeout as the migration's first statement, and a statement timeout on the serving unit.

Part 2 · Seeing what's running

Layers of visibility

A job's verdict belongs in Postgres, not the journal. An empty journal read is not evidence.handbook, platform notes

LayerSignalWhere it goesCatches
KernelPSI (/proc/pressure), cgroup memory.events, sar every 10 minmemguard's gate; incident forensicsPressure building before anything fails
systemdUnit result and exit statusOnFailure= → alert unit → Hermes → Slack, with the journal tail attachedA service or job that fails (114 of 114 first-party units wired)
systemdsystemctl --failed, swept at :09, :29, :49Repeat alert while anything stays failedA failure whose first alert was missed
journaldPer-unit identifier, persistent, 30 daysOperators and agents via journalctlWhat a unit said, and when
ApplicationScheduler heartbeat row, stale after 180 sTrader summary endpoint and consoleA loop that's alive but stuck
ApplicationCapture status file and a bars-written watchdogAPI, auto-halt and alertsA feed that's "connected" but silent
DataFeed-staleness check on the data itselfChecker registryA job that reports success while writing nothing
system-monitor15-minute push of host metrics, services, databases and probesController alerts after 2 consecutive failures, once per day per probeHosts and outcomes going wrong; a staleness job watches for silent hosts
Checker registry40 checks at 07:07 and 19:07One idempotent work item per red, never auto-closedConfiguration drift, missing alerts, overruns, undeployed code
Morning reviewOvernight job roster, day ledger, deploy driftA CLEAN / FOLLOW-UPS / PROBLEM reportAnything that slipped past the rest

Commands worth knowing

Read-only commands of the kind used to gather this page.

cat /proc/pressure/memory

Share of time tasks stalled on memory. "full" means everything runnable was waiting. The 6 September freeze was a PSI story long before it was an outage.

cat /sys/fs/cgroup/user.slice/user-1001.slice/memory.events

How often the seat hit MemoryHigh (throttled) or MemoryMax, and how many OOM kills happened inside it.

systemctl show user-1001.slice -p MemoryCurrent,MemoryHigh,MemoryMax,MemorySwapMax

The limits in force and current use, without reading cgroupfs by hand.

sar -r -f /var/log/sysstat/saDD -s 17:00:00 -e 19:30:00

Memory history for day DD from sysstat. This is how the 6 September timeline was rebuilt after the reset.

journalctl -u pg-backup-restore-smoke.service --since 2026-09-28 -o cat

One unit's messages, bare. Here it prints the weekly restore test passing.

journalctl -t nexus-unit-alert --since -30d

Filter by identifier across every instance of the alert template: 543 alerts for 17 units in 30 days.

journalctl --since -1d -p err

Everything logged at error priority or worse in the last day, host-wide.

systemctl show nexus-trader-scheduler -p ActiveEnterTimestamp,NRestarts,Result,ExecMainStatus

When it started, how often systemd restarted it, and how the last run ended.

systemctl --failed; systemctl --user --failed

The two failure lists: system services, and the dev user's CI and agent units.

systemd-analyze security nexus-trader-capture

A hardening score per unit. The capture scores 5.8; most unhardened services score 9.2.

Gaps, and what I'd change

Sources: read-only inspection of the production host on 2026-09-30 (cgroupfs, /proc/pressure, systemctl show and cat, journalctl, systemd-analyze); Postgres settings from pg_settings and hit ratios from pg_stat_database; the handbook's incident report for 6 September, platform notes and ops tooling; unit files on each repository's main branch.