One Linux server runs the market database, paper trading, ten APIs, backtests, 18 CI runners and a fleet of AI coding agents. The rule is simple: the database and trading keep running, whatever everything else does. This page shows how accounts, cgroups and scheduling enforce that, what happened when they didn't, and how the host is watched.
Postgres floor18 GiBon Postgres and both parent slices; no ceiling
Agents + CI seat34 / 36 GiBMemoryHigh / MemoryMax, 12 of 16 CPUs
Swap for agent jobs0an overrun is killed at once, never swapped
Journal30 dayspersistent, capped at 16 GB
Read from the production host on 2026-09-30, read-only: cgroupfs, /proc, systemctl and journalctl.
Part 1 · Separating the workload
Accounts: who runs what
Each service runs as its own Linux user with no login shell, owns its own directory, and keeps its secrets in a file only it can read. The agents' account can read everything and change nothing directly.
Account
Runs
Can
Cannot
market-data
Data API, analysts, data timers
Write its own directory under /opt; its .env is mode 0600
Log in; write system directories (ProtectSystem=full)
nexus-trader
Trader API, price capture, scheduler
Write only /var/lib/nexus-trader (ProtectSystem=strict)
Gain privileges (NoNewPrivileges)
nexus-sim-api
Research API and backtests
Run spawned worker pools
Take CPU from trading (CPUWeight=20)
7 more
pms, risk, arb, planner, RFQ simulator, monitor, research data
One service each
Log in
postgres
PostgreSQL 16 + TimescaleDB
Keep an 18 GiB memory floor, held all the way up the tree
Hit a memory ceiling: it has none, by design
dev
AI agents, CI runners, tooling
Read the system journal; read all database data (pg_read_all_data); queue deploys and jobs
sudo; change another service's files or secrets
admin
Administration by a person
sudo (the only account with it)
How dev triggers a privileged action without sudo
devdeploy-now traderwrites a request file into a queue directory (ACL: dev rwx)
root · systemd.path unit firesDirectoryNotEmpty= or PathExistsGlob=*.req
root · dispatcherChecks before actingflock -n · name must be on a root-owned allow-list · refuses if that list is group- or world-writable
service userRuns the jobrunuser -u <user> · env -i · timeout --kill-after=30 · args checked by regex, ≤ 4096 bytes
Only a name crosses the boundary. Everything that runs is decided by root-owned configuration, so a compromised agent account can queue a deploy but can't choose what it executes.
cgroups: who gets protected, capped or weighted
cgroup v2 with the cpu, memory and pids controllers enabled. Services sit in system.slice; everything the agents and CI do sits under the dev user's slice, called the seat below, which carries the limits.
protect memory floor limit ceiling or quota weight share under contention
/controllers: cpu memory pids · io not enabled
system.sliceMemoryMin=18Gno ceiling, on purpose
nexus-*.service~24 long-running services, each as its own user
nexus-sim-backtest.serviceCPUWeight=20about 6% vs 25% of CPU under contention; every core when idle
pg-backup-offsite.serviceCPUWeight=10High 1G · Max 2Gswap 0
RAM to scale. The seat's ceiling is what remains after the floors: 61.9 − 18 − 5 − 2 − 1 ≈ 36 GiB. MemoryHigh sits 2 GiB below MemoryMax, so a job that drifts over gets throttled before it gets killed. Before 6 September the caps added up to more than the machine had.
Why 18 GiB for Postgres
The floor is the database's buffer pool plus a little headroom. It sounds generous until you see what it does and doesn't cost.
shared_buffers16 GiBabout 25% of RAM, the usual sizing
Headroom+2 GiBconnection processes
In use now19.2 GiB16.2 of it the buffer pool
Buffer hit ratio99.9%market_data, since stats began
It protects; it doesn't reserve.memory.min tells the kernel not to reclaim Postgres's memory below 18 GiB when memory runs short. Any part of the floor Postgres isn't using is free for everything else, and today it already uses more than the floor.
It can't sensibly go below 16 GiB. The buffer pool lives in shared memory, which the kernel can't drop the way it drops page cache; it can only swap it out. A floor smaller than the pool lets part of it go to swap, and then the reads that hit memory 99.9% of the time wait on disk instead. That is how the database became unreachable on 6 September.
The 2 GiB of headroom is modest. Idle connections use almost nothing. The worst case is far larger: 32 MB of work_mem per sort across up to 100 connections, plus 2 GiB of maintenance_work_mem for each vacuum, index build or compression job. The floor protects the working set; it doesn't try to bound that.
The budget was built around it. Floors total 26 of 62 GiB: 18 for Postgres and 8 for CI. The agents' 36 GiB ceiling was worked out after subtracting this one.
To shrink it, change the pool. An 8 GiB pool with a floor near 10 GiB would lean more on the operating system's cache and give the agents more room. The cost would be more disk reads on a 130 GB database with hot recent data, measured by the hit ratio and query latency, and it needs a database restart.
Found and fixed while writing this page · 30 September 2026
A floor only counts if every parent has one
In cgroup v2 the kernel scales each level's protection by its parent's effective floor, and a parent at zero gives zero. Postgres has no memory ceiling of its own, so the only reclaim that reaches it is host-wide, the 6 September condition. Its two parent slices had no floor, so its 18 GiB was protecting nothing in exactly that case. The existing check read only Postgres's own value and stayed green.
cgroup
memory.min before
after
system.slice
0
18 GiB
system-postgresql.slice
0
18 GiB
postgresql@16-main
18 GiB
18 GiB
Effective under host-wide reclaim
0
18 GiB
Root added a MemoryMin=18G drop-in to both slices at about 09:40 UTC, reloaded systemd and wrote the live values, without restarting the database. The remaining work is a check that reads every ancestor, with a negative control, so this can't pass silently again.
The other levers
In force
CPUWeight, not nice
CPUWeight=20 # backtests
CPUQuota=1200% # dev seat
Under cgroup v2, nice only orders processes inside one cgroup, so it can't protect one service from another. Weights share CPU only under contention, so backtests still use every core on a quiet box.
Each heavy job gets its own cgroup with swap off. A test allocation over the cap was killed with rc 137 and zero swap used. It warns when the request exceeds the seat's headroom.
Gate on · freezer observing
A memory gate on dispatch
LOADED if headroom < 10 GiB,
MemAvailable < 10 GiB or
PSI memory some avg60 ≥ 10%
One reader, memguard, decides whether new agent work may start. Dispatch refuses (exit 18) when closed. It can freeze running jobs with the cgroup freezer rather than kill them, but today it only logs. It covers the ground systemd-oomd would, which isn't installed; freezing rather than killing suits long agent jobs that can resume.
A one-minute timer asks the kernel to drop the seat's page cache (up to 4 GiB per tick), so cache can't push the seat into throttling.
In force
/tmp is RAM
/tmp: 31G tmpfs
TMPDIR → on-disk path
tmp-reaper every 6 h
Files in a tmpfs count as shared memory. Dead session directories and orphaned test databases once filled 24 GiB of it, so work now writes to disk and a reaper clears leftovers.
Root dispatchers skip if another pass holds the lock. Every job has a hard timeout. Migrations give up waiting for a lock rather than queueing everything behind them.
Trading and the broker rate limiter run on the quarter-hours, so other daily jobs avoid them. The rule came from a login collision on 31 August.
In force
Task limits
TasksMax=33% # dev seat, 153,710
TasksMax=256 # job runner
A fork bomb or thread leak in the seat stops at its own limit. The seat has never reached it.
Not available
I/O isolation
io controller: not enabled
NVMe scheduler: none
IOWeight needs the io controller and ionice needs the BFQ scheduler; neither is in place. Disk contention is managed by memory floors that keep hot data cached instead.
When the host stopped responding
Each episode left a control behind. The timelines come from sar history (sysstat samples every 10 minutes), the persistent journal and cgroup event counters.
Swapped solid for 100 minutes
Three agent sessions ran mutation tests and noise experiments at once. Memory rose by about 21 GiB an hour. The seat had a MemoryMax but unlimited swap, and MemoryMax bounds RAM only, so the kernel swapped instead of killing anything. No OOM kill appears in that boot's log. Postgres had no floor, lost its cache and went to disk. ssh answered but couldn't fork a login, and the box needed a hardware reset.
Added: swap caps, capped scopes for heavy jobs, the 18 GiB Postgres floor, seat limits recomputed from the RAM budget, and a memory alert that was proven to deliver.
Page cache slowed CI
The seat sat at 33.2 of 34 GiB, mostly cache, and was throttled millions of times. A 40-minute CI run took over an hour.
Added: the page-cache reclaim timer.
/tmp full of memory
24 GiB of dead session directories, test output and 26 orphaned test Postgres servers sat in the RAM-backed /tmp. The seat cap held, so the database kept running, but ssh froze for about 15 minutes.
Added: an on-disk TMPDIR and the reaper.
Dispatches piled on
The seat's working set rose from 19 to 32 GiB in 15 minutes while new agent work kept arriving. ssh stalled and six jobs were frozen by hand.
Added: memguard, one memory verdict that dispatch, claims and drains all read.
The database, not the host
A deploy's migration waited for a lock behind a 57-minute query, and 45 connections queued behind the migration. The API was down for 33 minutes while the host was fine.
Added:lock_timeout as the migration's first statement, and a statement timeout on the serving unit.
Part 2 · Seeing what's running
Layers of visibility
A job's verdict belongs in Postgres, not the journal. An empty journal read is not evidence.handbook, platform notes
Layer
Signal
Where it goes
Catches
Kernel
PSI (/proc/pressure), cgroup memory.events, sar every 10 min
memguard's gate; incident forensics
Pressure building before anything fails
systemd
Unit result and exit status
OnFailure= → alert unit → Hermes → Slack, with the journal tail attached
A service or job that fails (114 of 114 first-party units wired)
systemd
systemctl --failed, swept at :09, :29, :49
Repeat alert while anything stays failed
A failure whose first alert was missed
journald
Per-unit identifier, persistent, 30 days
Operators and agents via journalctl
What a unit said, and when
Application
Scheduler heartbeat row, stale after 180 s
Trader summary endpoint and console
A loop that's alive but stuck
Application
Capture status file and a bars-written watchdog
API, auto-halt and alerts
A feed that's "connected" but silent
Data
Feed-staleness check on the data itself
Checker registry
A job that reports success while writing nothing
system-monitor
15-minute push of host metrics, services, databases and probes
Controller alerts after 2 consecutive failures, once per day per probe
Hosts and outcomes going wrong; a staleness job watches for silent hosts
Checker registry
40 checks at 07:07 and 19:07
One idempotent work item per red, never auto-closed
Read-only commands of the kind used to gather this page.
cat /proc/pressure/memory
Share of time tasks stalled on memory. "full" means everything runnable was waiting. The 6 September freeze was a PSI story long before it was an outage.
One unit's messages, bare. Here it prints the weekly restore test passing.
journalctl -t nexus-unit-alert --since -30d
Filter by identifier across every instance of the alert template: 543 alerts for 17 units in 30 days.
journalctl --since -1d -p err
Everything logged at error priority or worse in the last day, host-wide.
systemctl show nexus-trader-scheduler -p ActiveEnterTimestamp,NRestarts,Result,ExecMainStatus
When it started, how often systemd restarted it, and how the last run ended.
systemctl --failed; systemctl --user --failed
The two failure lists: system services, and the dev user's CI and agent units.
systemd-analyze security nexus-trader-capture
A hardening score per unit. The capture scores 5.8; most unhardened services score 9.2.
Gaps, and what I'd change
The Postgres floor check reads one levelit passed while the floor was ineffective
The slice floors are live. The check still needs to read every ancestor, with a negative control that must go red on a zero parent. That work is scheduled.
The memory governor only observesmemguard's freezer has run in observe mode since 25 Sep
The dispatch gate already refuses new work; it read CLOSED on the morning this page was written. Switching the freezer to enforce would also pause running agent jobs before the seat stalls.
No I/O isolationio controller not enabled
Delegating the io controller and setting IOWeight would stop a heavy backfill from competing with database reads.
Hung processes can look healthyno WatchdogSec= on any first-party unit
A process that hangs but stays alive is only caught where an app heartbeat exists. The trader loop and the capture should use Type=notify with a systemd watchdog.
The dev user's units don't page113 user units; the alert handler is system-only
A twice-daily check covers them. A user-level alert path with its own credential would close the gap.
Most services are unhardened9.2 on systemd-analyze security
Copying the trader's hardening block to the research, portfolio and arbitrage services is mechanical work.
Sources: read-only inspection of the production host on 2026-09-30 (cgroupfs, /proc/pressure, systemctl show and cat, journalctl, systemd-analyze); Postgres settings from pg_settings and hit ratios from pg_stat_database; the handbook's incident report for 6 September, platform notes and ops tooling; unit files on each repository's main branch.