deploy/verify.sh
Trusts CI, and drops to smoke tests, only if a job named exactly pytest-full passed on the commit being deployed, or on an identical file tree. So no reduced lane may ever produce that name.
Nexus explainer · companion to Stack by example
Nexus runs its CI on its own GitHub Actions runners. Most of the design answers one question: how much of a 142-minute test suite does a given change have to pay for? This page shows how a change is classified, where the tests run, what a green check means and how code reaches main.
Read from the GitHub API and each repository's main branch on 2026-09-30. Run totals cover 16–30 September; job figures come from the 600 most recent research-suite runs, 26–30 September.
Every pull request to the research service passes through a classifier that decides how much of the suite the diff needs: a docs check for a docs change, a few minutes for a change inside one package, the whole suite for anything uncertain. The rules that pick the lane are read from the trusted base commit, so a pull request can't edit them to lower its own bar, and every failure path falls back to the full suite.
Merged code reaches main through one writer, a batch lane that merges tested groups of pull requests. When a merge's exact file tree has already passed the full suite, the push to main skips the suite.
The research service has the largest test suite by far, so almost all CI time is spent there. Inside that suite, 93% of job time is the pytest partitions themselves; classification, lookups and bookkeeping are the rest.
| Repository | Runs | Wall-clock hours | Share |
|---|---|---|---|
| nexus-sim-apiresearch service | 3,659 | 772.7 | 76.5% |
| nexus-engineering-handbookoperations tooling and guards | 1,047 | 168.0 | 16.6% |
| nexus-data-apimarket-data service | 471 | 56.8 | 5.6% |
| 11 other repositories | 639 | 12.9 | 1.3% |
Workflow runs created 16–30 September 2026. Wall clock is how long people waited, not what the machines did: it includes queueing, and 19% of it is runs that were later cancelled.
A classify job reads the diff and picks a lane. Each selector it runs is taken from the base commit with git show "${BASE_SHA}:scripts/…", never from the pull request's own code. An unresolvable diff, an unreadable selector or an unknown result all fall to full.
| Lane | Chosen when | What runs | Median job-minutes | Runs |
|---|---|---|---|---|
| full | Anything else, and every uncertainty: unmapped source, broad script or data paths, the nightly run, a main push without attestation. | The whole suite, about 7,400 tests | 53.3 | 119 |
| phase2-skip | Every changed module lies provably outside the import closure of the five heavyweight proof files. | The suite without those proofs | 32.6 | 115 |
| scoped | A package map covers the change: 22 source packages mapped to their tests. | That package's tests plus 8 canaries kept on any source change | 5.7 | 65 |
| api | Import-closure analysis reaches only API code. It raises rather than guess on a dynamic import, eval or star-import. | A narrow API selection | — | 0 |
| ci-config | The diff touches only CI configuration on an allow-list, and not the policy that defines the allow-list. | The CI unit tests | — | 0 |
| docs | Markdown only, or Python edits whose syntax tree is unchanged. Refused when a doc is itself a test input. | A docs check, no pytest | 1.3 | 17 |
| tree-attested | A push to main whose exact file tree already passed a full run on its pull request. | Nothing: the earlier result stands | 1.8 | 40 |
Runs are the 600 most recent research-suite runs, 26–30 September; job-minutes are summed over every job in a successful run. Two more paths reach the required check without running a suite: batch members, whose suite runs once for the whole batch (143 runs, 0.7 job-minutes), and label-only events, which reuse the head commit's recorded result (93 runs, 0.9 job-minutes).
Five proof files hold 33 tests and 89.2 of the suite's 141.9 test-minutes, 63% of the total. That is why a lane exists just to prove a change can't reach them.
Every partition job pays the same fixed cost: checkout, a locked install, lock-parity and engine-pin checks, an artifact upload. A scoped run still holds three runners, just not for long.
Of 524 pull-request runs, 16% ran the full suite. Batch deferral, the phase-2 skip, result reuse and scoping covered almost all the rest.
Whatever the lane, the selection runs as a three-way matrix. All three partitions collect the same selection and filter it locally. An evidence job then downloads every partition's manifest and proves they form an exact disjoint union of what was collected, with matching commit, tree, engine, Python and lock hash, all from the same run attempt.
| Partition | Runs on | Workers | Holds | Median min, full |
|---|---|---|---|---|
| portable-1 | compute host, any of 3 runners | 6 | shard 1 of the portable map: 14 named groups | 27.9 |
| portable-2 | compute host, any of 3 runners | 6 | shard 2: 10 named groups | 20.2 |
| host | production host, 2 runners | 2 | exactly five host-bound files | 3.2 |
Host-bound is a contractnot a cost decision
Two files pin exact BLAS decision digests and need the production CPU; two check the installed deploy tooling; one observes the live database. Everything else is portable and can run on any compute runner.
A committed shard mapci-portable-shards.json
24 named test groups totalling 181.8 CPU-minutes are assigned whole, 36 heavy single tests are pinned, and everything left over goes by sha256(nodeid) % 2. The split is decided in the repository, not by the runner.
More workers is not faster6 wide beat 12 wide
On the old compute host, twelve workers took 52.9 minutes against six workers' 45. The worker budget caps a full portable run at six and gives the host partition two.
Re-run all of it or nonesame-attempt rule
Re-running one failed partition leaves the other two on the previous attempt. The verifier then finds no matching manifests and the check goes red with every test passing.
25 self-hosted runners serve 14 repositories: 18 on the production host and 7 on a separate compute host. Each runs as a systemd user service in its own slice. Workflows ask for a capability label, never a machine.
compute host · 3 research runners
The portable partitions, 20 to 30 minutes each. When the old compute host was retired on 22 September, its work moved by registering new runners with this label.
production host · 2 runners
The host partition, classify, the tree lookup, the docs check, secret scanning, the required pytest aggregate and batch bookkeeping. Every job here takes seconds or a few minutes.
production host · 14 runners · compute host · 4
Every other repository has its own runner on the production host for its whole workflow, two for the handbook. The handbook and the market-data service also run their heavier lanes on the compute host.
No label set has one runnera missing runner doesn't fail, it waits
A job whose labels match no online runner is not an error: it queues indefinitely, and the check simply never arrives. In the research service, every label set is served by at least two runners.
The reserved slothost runners never take compute work
If a host-partition runner also carried the compute label, 20-to-30-minute portable jobs could occupy every runner the host partition can use. That once left the host partition queued for over three hours.
Queueing is raremedian wait ≈ 0 min
Once a partition job is created it usually starts at once. The 90th-percentile wait is 8 to 13 minutes for portable jobs and under a minute for the host partition.
The required checks on main are pytest and secret-scan, deliberately fixed names, because the suite job's own name is computed from the lane at run time and a dynamic name can't be a required check. Everything else keys off exact job names, in four places that share no code.
Trusts CI, and drops to smoke tests, only if a job named exactly pytest-full passed on the commit being deployed, or on an identical file tree. So no reduced lane may ever produce that name.
In a different repository. A batch is full only with pytest, secret-scan, pytest-full, the host partition and both portable shards. One shard alone reads as not full.
One manifest per worker; identical scope, canaries and collected selection across the three partitions; all from the same run attempt.
Defines "this tree has been tested": the squash tree that already passed pytest-full on a pull request, which lets the push to main skip the suite.
A change to the shape of CI can break consumers its own pull request never exercises. When the portable suite was first split into two shards, the change was green on its own pull request and then broke four things in turn, the last of them in the batch lane's repository. A pull request only ever tests the lane it happens to take.
The research service's main branch is protected and has one writer: a batch lane that runs as two systemd timers in the handbook repository, not as a workflow. It reads the research CI's job names as evidence. Arming GitHub's auto-merge on a pull request would give main a second writer, so the lane refuses members that have it on.
| How it merged | Tick | Admits | Push to main, 16–30 Sep |
|---|---|---|---|
| FULL batch | 60 s | members whose own lane was full | tree-attested 95 of 106 pushes, median 1.9 min |
| FAST batch | 30 s | members whose own lane was docs, scoped or api | full suite 32 of 33 pushes, median 37.6 min |
| Merged by hand | — | outside the lane | full suite 33 of 37 pushes, median 48.0 min |
Admission fails closedevery member, every time
A member needs a batch label, a pull-request body that opens with its merge class and author, and an approval from a different agent that names the exact 40-character head commit. Classes are tooling, test and docs; engine changes need a separate, operator-granted label. At most 16 members, in pull-request order.
A batch can't test its own CICI changes are refused
Any member touching .github/workflows/ or the CI scripts is refused whatever its class, so a batch is never judged by CI it modified.
The two lanes differ by designattested versus full
A FULL batch's members each produced a pytest-full, so its squash tree is already attested and the push to main costs two minutes. A FAST batch never produces one, so its push runs the whole suite.
A hand-merge costs everyoneit moves the batch base
On the day the lane went live, a docs pull request merged by hand moved the base under a green batch, retired it and cost a 50-minute rerun. Since then main has one writer.
Coverage floorsnexus-trader
The only repository that gates coverage: a 68.41% total floor and 14 per-module floors, each set five points below the module's measured coverage.
Anti-vacuity floorsdata API, janus
A run that collects too few tests fails: under 5,000 in the data API's portable lane, 85 in its host lane, 30 in janus. A suite that silently stops collecting can't pass.
A lint ratchetnexus-data-web
The lint job fails only when errors rise above a recorded baseline of 63, so the count can only go down.
Schedulesthree, estate-wide
A full nightly run of the research suite at 03:30 UTC, a daily pip-audit --strict against the lockfile in 8 repositories, and one six-hourly estate check.
Deliberately absentno shared actions
No reusable workflows, composite actions or path filters. Shared files are copied: the secret-scan workflow is byte-identical in 8 repositories.
Sources: the GitHub REST API for workflow runs (16–30 September 2026) and jobs (the 600 most recent research-suite runs, 26–30 September), runner registrations and branch protection; workflow files and CI scripts on each repository's main branch, read on 2026-09-30. Merge-push figures classify every push to main in the window by the branch it came from.