London Data

Nexus explainer · companion to Stack by example

The Nexus CI System

Nexus runs its CI on its own GitHub Actions runners. Most of the design answers one question: how much of a 142-minute test suite does a given change have to pay for? This page shows how a change is classified, where the tests run, what a green check means and how code reaches main.

Workflow runs5,816in 14 days, 1,010 wall-clock hours, 14 repositories
One workflow73%of those hours: the research service's test suite
Self-hosted runners2518 on the production host, 7 on a compute host
Full run30 minmedian wall clock; 53 job-minutes over three partitions
Scoped run6 minmedian wall clock for a change inside one package
GitHub-hosted jobs0in the research suite since 22 September

Read from the GitHub API and each repository's main branch on 2026-09-30. Run totals cover 16–30 September; job figures come from the 600 most recent research-suite runs, 26–30 September.

In short

Every pull request to the research service passes through a classifier that decides how much of the suite the diff needs: a docs check for a docs change, a few minutes for a change inside one package, the whole suite for anything uncertain. The rules that pick the lane are read from the trusted base commit, so a pull request can't edit them to lower its own bar, and every failure path falls back to the full suite.

Merged code reaches main through one writer, a batch lane that merges tested groups of pull requests. When a merge's exact file tree has already passed the full suite, the push to main skips the suite.

One change, end to end

Pull request diff against base classify rules from base commit Lane docs … full portable-1 · compute portable-2 · compute host · production Evidence exact disjoint union Required checks pytest · secret-scan Batch lane one writer to main Push to main attested or full Deploy gate trusts pytest-full both green Everything on the upper row runs for every pull request; the lower row runs once a pull request is approved and batched.
The two amber boxes are the decisions: which lane a diff takes, and whether the partitions together really ran what was selected. Every downstream consumer reads job names, so those names are part of the contract.

One repository is the bill

The research service has the largest test suite by far, so almost all CI time is spent there. Inside that suite, 93% of job time is the pytest partitions themselves; classification, lookups and bookkeeping are the rest.

RepositoryRunsWall-clock hoursShare
nexus-sim-apiresearch service3,659772.776.5%
nexus-engineering-handbookoperations tooling and guards1,047168.016.6%
nexus-data-apimarket-data service47156.85.6%
11 other repositories63912.91.3%

Workflow runs created 16–30 September 2026. Wall clock is how long people waited, not what the machines did: it includes queueing, and 19% of it is runs that were later cancelled.

Seven lanes, chosen from the base commit

A classify job reads the diff and picks a lane. Each selector it runs is taken from the base commit with git show "${BASE_SHA}:scripts/…", never from the pull request's own code. An unresolvable diff, an unreadable selector or an unknown result all fall to full.

LaneChosen whenWhat runsMedian job-minutesRuns
fullAnything else, and every uncertainty: unmapped source, broad script or data paths, the nightly run, a main push without attestation.The whole suite, about 7,400 tests53.3119
phase2-skipEvery changed module lies provably outside the import closure of the five heavyweight proof files.The suite without those proofs32.6115
scopedA package map covers the change: 22 source packages mapped to their tests.That package's tests plus 8 canaries kept on any source change5.765
apiImport-closure analysis reaches only API code. It raises rather than guess on a dynamic import, eval or star-import.A narrow API selection—0
ci-configThe diff touches only CI configuration on an allow-list, and not the policy that defines the allow-list.The CI unit tests—0
docsMarkdown only, or Python edits whose syntax tree is unchanged. Refused when a doc is itself a test input.A docs check, no pytest1.317
tree-attestedA push to main whose exact file tree already passed a full run on its pull request.Nothing: the earlier result stands1.840

Runs are the 600 most recent research-suite runs, 26–30 September; job-minutes are summed over every job in a successful run. Two more paths reach the required check without running a suite: batch members, whose suite runs once for the whole batch (143 runs, 0.7 job-minutes), and label-only events, which reuse the head commit's recorded result (93 runs, 0.9 job-minutes).

Where the minutes sit

Five proof files hold 33 tests and 89.2 of the suite's 141.9 test-minutes, 63% of the total. That is why a lane exists just to prove a change can't reach them.

Cheap in wall time, not in occupancy

Every partition job pays the same fixed cost: checkout, a locked install, lock-parity and engine-pin checks, an artifact upload. A scoped run still holds three runners, just not for long.

What pull requests actually take

Of 524 pull-request runs, 16% ran the full suite. Batch deferral, the phase-2 skip, result reuse and scoping covered almost all the rest.

Inside a run: three partitions, reconciled afterwards

Whatever the lane, the selection runs as a three-way matrix. All three partitions collect the same selection and filter it locally. An evidence job then downloads every partition's manifest and proves they form an exact disjoint union of what was collected, with matching commit, tree, engine, Python and lock hash, all from the same run attempt.

PartitionRuns onWorkersHoldsMedian min, full
portable-1compute host, any of 3 runners6shard 1 of the portable map: 14 named groups27.9
portable-2compute host, any of 3 runners6shard 2: 10 named groups20.2
hostproduction host, 2 runners2exactly five host-bound files3.2

Runners are chosen by capability, not by host

25 self-hosted runners serve 14 repositories: 18 on the production host and 7 on a separate compute host. Each runs as a systemd user service in its own slice. Workflows ask for a capability label, never a machine.

nexus-compute

compute host · 3 research runners

The portable partitions, 20 to 30 minutes each. When the old compute host was retired on 22 September, its work moved by registering new runners with this label.

nexus-host · nexus-docs

production host · 2 runners

The host partition, classify, the tree lookup, the docs check, secret scanning, the required pytest aggregate and batch bookkeeping. Every job here takes seconds or a few minutes.

Everything else

production host · 14 runners · compute host · 4

Every other repository has its own runner on the production host for its whole workflow, two for the handbook. The handbook and the market-data service also run their heavier lanes on the compute host.

What a green check means

The required checks on main are pytest and secret-scan, deliberately fixed names, because the suite job's own name is computed from the lane at run time and a dynamic name can't be a required check. Everything else keys off exact job names, in four places that share no code.

Deploy gate

deploy/verify.sh

Trusts CI, and drops to smoke tests, only if a job named exactly pytest-full passed on the commit being deployed, or on an identical file tree. So no reduced lane may ever produce that name.

Batch lane

merge_batch/policy.py

In a different repository. A batch is full only with pytest, secret-scan, pytest-full, the host partition and both portable shards. One shard alone reads as not full.

Evidence

verify_ci_partitions.py

One manifest per worker; identical scope, canaries and collected selection across the three partitions; all from the same run attempt.

Attestation

tree_attestation.py

Defines "this tree has been tested": the squash tree that already passed pytest-full on a pull request, which lets the push to main skip the suite.

A change to the shape of CI can break consumers its own pull request never exercises. When the portable suite was first split into two shards, the change was green on its own pull request and then broke four things in turn, the last of them in the batch lane's repository. A pull request only ever tests the lane it happens to take.

Main has one writer, and it is a timer

The research service's main branch is protected and has one writer: a batch lane that runs as two systemd timers in the handbook repository, not as a workflow. It reads the research CI's job names as evidence. Arming GitHub's auto-merge on a pull request would give main a second writer, so the lane refuses members that have it on.

How it mergedTickAdmitsPush to main, 16–30 Sep
FULL batch60 smembers whose own lane was fulltree-attested 95 of 106 pushes, median 1.9 min
FAST batch30 smembers whose own lane was docs, scoped or apifull suite 32 of 33 pushes, median 37.6 min
Merged by hand—outside the lanefull suite 33 of 37 pushes, median 48.0 min

The other repositories

Lessons written down along the way

What I would change

Sources: the GitHub REST API for workflow runs (16–30 September 2026) and jobs (the 600 most recent research-suite runs, 26–30 September), runner registrations and branch protection; workflow files and CI scripts on each repository's main branch, read on 2026-09-30. Merge-push figures classify every push to main in the window by the branch it came from.