Nexus case study · companion to the full case study
Nexus is built and maintained by AI coding agents from three vendors, directed by one engineer. This page shows how that works day to day: the seats, the Postgres work board that holds every piece of work, the lease a seat takes before it edits anything, the dispatcher that hands work out, and the review, merge and deploy steps a change goes through before it reaches production.
Read from the work board, its claim ledger and the dispatcher's audit log on 2026-10-01 at 19:58 UTC.
Work is a row in Postgres, not a chat message. A seat takes a row with an atomic lease and gets its own git worktree. It opens a pull request, a different agent approves the exact commit, CI runs the tests the diff needs, a merge lane lands it, and a deploy tool ships it and writes a ledger row. When the lease ends, the seat records how it ended.
Handing work to a seat, taking a lease, merging and deploying are gated in code, not left to an agent's judgement. A gate that can't read what it needs refuses. The rules are written down once, in a handbook that every vendor's agents install from the same branch.
take decides who owns it, and review is done by a different agent. The dashed lines are how work gets back onto the board.A seat is a long-running agent session in its own terminal. A seat exists only if it has a line in the roster file in the handbook. Four documents once disagreed about how many seats there were (19, 25, 25 and 26), so now the roster is the only place a seat is defined, and every tool reads it from there. Three more seats on a second machine use the same board.
Each tile is one seat, coloured by vendor.
Roles live in the rosternot in what a seat says about itself
One seat coordinates: it turns plans into rows and hands them out. Reserved seats, including the operator's own, are never handed work by the dispatch tools.
Effort is fixed at launchon all three vendors
A row carries the reasoning tier it needs. The tier a seat runs at can only be changed by relaunching it, so a relaunch tool waits until the seat is idle and holds no lease.
Seats are not capacityquota is
An hourly meter records each vendor's remaining plan quota. A seat from a vendor that is out of quota is idle for that reason, and handing it work only burns the rest of the week.
Every piece of work is a row in one table: features, fixes, reviews, research tasks, questions for the operator and jobs that need root. A seat that has just been cleared knows nothing except the row id it was handed, so the row's body has to be enough to act on. If the brief isn't written in the row, the row isn't ready to hand out.
| Field | What it holds | Why it is there |
|---|---|---|
| Bodythe brief | The repository, the change and how to tell it's done, written so a fresh session can act on it. | The dispatcher won't hand out a row whose brief is too thin to act on. |
| Size, effort, riskS–XL · low–max · trivial–critical | How big the job is, the reasoning tier it needs, and what goes wrong if it's done badly. | A routine row worked at the top tier wastes quota; a hard one worked at the bottom tier gets a confident wrong answer. |
| Kindtask · watch · gate · finding | Work to do, a condition to keep checking, a date or event to wait for, or a finding to record. | 3,004 tasks, 35 watches, 15 gates and 2 findings. |
| Owner and laneagent · operator · root | Whether an agent can do the work, or it needs the operator's decision or a root command. | An agent blocked on a person or on root files the one command it needs as its own row. |
| Parent and campaign691 child rows · 56 campaigns | How work breaks down into smaller rows, and which programme it belongs to. | A campaign can be rebuilt from the board alone, after any agent's context has been cleared. |
| Notes and PRson the row itself | Short findings, and the pull requests that carry the work. | A seat's report goes on the row, so it outlives the terminal it was written in. |
inbox→ready→in_review→done·orblockedwith a reason ·cancelled
On 1 October 2026 the board held 2,816 done rows, 128 cancelled, 45 in the inbox, 39 blocked, 26 in review and 2 ready. Rows span 17 repositories. A row is created only after work-task ls has been checked for an existing one, because most work that looks missing already has a row.
Before its first edit, a seat runs work-task take. That one command claims the rows, records the lease, and creates a git worktree on a fresh branch from main. Seats never edit the shared clones. The same instruction can go to every idle seat at once: the take is an atomic compare-and-swap, so exactly one seat wins each row and the rest are refused.
Refused is a successexit 1: a sibling won
A seat that loses the race doesn't retry or force the claim. It picks a different row, or stops if nothing is free.
Leases expire on readno sweeper process
A lease has a fixed term, usually two hours. Any reader treats an expired lease as free, so a dead seat holds nothing for long. A seat extends its lease at about half-life; 829 extensions are on record.
Release says how it endedwith evidence
For a code change, done needs a merged pull request; while the pull request is open the work is in_review, and the release names it. Work with no pull request, such as research or a ruling, closes on other recorded evidence. The release runs as its own command, after the seat has read the output of its work.
One identity, one claimcollisions are refused
Two shells once reported the same agent id, and one released the other's live lease with no trace. A second claim under a shared id is now refused, and every forced release is logged: 64 so far.
Every lease taken, by the week it started. Hover or focus a segment for the count.
The mix moves with quota. When one vendor's weekly allowance runs low, work goes to the vendors that still have headroom.
| Week of | Claude | Codex | Grok | Operator | Total |
|---|---|---|---|---|---|
| 17 Aug (23 Aug only) | 108 | 0 | 351 | 0 | 459 |
| 24 Aug | 602 | 334 | 106 | 27 | 1,069 |
| 31 Aug | 148 | 257 | 66 | 0 | 471 |
| 7 Sep | 109 | 80 | 48 | 0 | 237 |
| 14 Sep | 202 | 192 | 373 | 6 | 773 |
| 21 Sep | 763 | 62 | 146 | 5 | 976 |
| 28 Sep (to 1 Oct) | 177 | 97 | 129 | 20 | 423 |
| Total | 2,109 | 1,022 | 1,219 | 58 | 4,408 |
The outcome recorded for each task released: 4,423 task releases from 4,017 leases. One lease can cover several tasks, so these are not counts of leases.
Short leasesmedian 6 min · p90 44 min, over 4,347 released agent leases
Rows are kept small. A seat takes at most three related rows at a time, so a large row doesn't starve the others.
Lapsed, not lost52 leases
A lease that ran out without a release is visible on the board, and the row goes back to being free.
Work reaches a seat through one dispatcher. It clears the seat's context, checks that the clear was accepted, then sends the row id. Most of its eleven checks run before any seat is touched, and a dispatch they refuse changes nothing. The rest run seat by seat and can stop the work after a seat has been cleared. Whether a clear really reset the session can be confirmed on some vendors' terminals, not all. The tools that relaunch seats share several of these checks.
| Check | Stops | Before any seat | Per seat |
|---|---|---|---|
| Memory | Any dispatch while the host is under memory load, or while its memory can't be read. | 7 | — |
| Budget | A vendor without enough quota left, or whose quota reading can't be trusted. | 3 | — |
| Standalone | An instruction that would leave a freshly cleared seat without enough to act on. | 3 | — |
| Clear | Sending work to a seat whose clear did not go through. | — | 3 |
| Effort | Work sent without settling the reasoning tier its row asks for. | 14 | 1 |
| Dispatchable | A row whose body is too thin to be a brief. | 45 | — |
| Routing | A vendor that an operator policy on the board excludes for this row. | 5 | — |
| Reserved | The coordinator's seat, the operator's seat and any other reserved seat. | 7 | — |
| Held | A seat that still holds a live lease, since clearing it would orphan the work. | 12 | — |
| New task under continue | New work sent to a seat without clearing it first. | not logged | — |
| Review envelope | A review request that doesn't ask the reviewer to approve a named commit. | not logged | — |
| Unsent textper seat, outside the eleven | A seat whose input box already holds text someone typed. Nothing is typed or cleared. | — | 5 |
| Dialog on screenper seat, outside the eleven | A seat showing a prompt that a keypress would answer. Nothing is typed. | — | 1 |
Counts are from the dispatcher's audit log of live dispatches, 22 September to 1 October. "Before any seat" refusals stopped the whole dispatch. "Per seat" stops left the work unsent on that seat; for the clear and effort stops, the seat may already have been cleared. The two checks marked "not logged" don't write an audit row when they refuse.
Prompts on screen are respectedone sender per seat
A keypress that lands on a prompt can answer it, and two seats were once lost that way. The dispatcher looks at each seat's screen before it sends anything, leaves a seat alone while a prompt is showing, and lets only one sender work on a seat at a time.
Sent is not submittednever prints a tick
Of 1,841 sends, the dispatcher saw 1,144 start work. It classed 648 as not yet started when it looked, a reading that includes known misreads of the screen, and couldn't tell for 49. It reports each one as it saw it and claims nothing more, and the coordinator follows up the rest.
A false idle costs a sessionso ambiguity refuses
The tool that relaunches idle seats weighs two errors. Wrongly reading a seat as busy costs one missed relaunch. Wrongly reading it as idle destroys a live session and whatever is uncommitted. Every signal it can't read counts as busy.
All the agents push to GitHub under one account, so GitHub's own review approvals can't tell them apart. An approval is a pull-request comment instead, naming the reviewing agent, the verdict and the full 40-character commit it read. An approval of any other commit doesn't count.
Reviews go to a seat other than the author's, usually from another vendor. Findings go back to the author through the board, and the review is repeated on the new head.
An author never turns on auto-merge for its own pull request. That happens only after an approval of the exact head commit. The research service has one writer to main, its batch lane, which rejects pull requests that have auto-merge on.
Merging isn't shipping. deploy-now ships the tested tree, runs the service's gate and a health check, and writes a row to the deploy ledger. It refuses dirty trees, stale clones and rollbacks; fixes go forward.
How each change gets only the tests it needs, and how the batch lane works: The Nexus CI system.
An agent's context is cleared between tasks, and a cleared session remembers nothing. So anything worth keeping has to live somewhere that survives a clear and can be read by any vendor's agents.
| It is | It lives in | Shared with |
|---|---|---|
| Work, findings, blockers | A row, its notes and its pull requests on the board | Every seat and the operator, queryable |
| Procedures | Nine skills in the handbook: claiming work, carrying it to a pull request, checking it, wrapping up, overnight review, and more | All three vendors, installed from the handbook's main branch |
| Rules that bind every seat | One page in the handbook, linked from every skill rather than copied into it | Every seat. A copy doesn't change when the rule does |
| Dates, gates and durable facts | The handbook's status page | Every seat, through git |
| Facts about the system | The code, the database or the deploy ledger, read when needed | Never written down as prose, because prose like "live is commit X" goes stale |
| Agent memory | A small store on the host | Hints only. If another agent needs to know it, it belongs on the board, in the handbook or in code |
The coordinator clears toothe board carries the campaign
One coordinating session ran 9,636 turns across six leases because it kept a whole campaign in its context. Now the coordinator writes every next step to the board before a clear, and checks that it can rebuild the campaign from two board commands alone.
The host comes firstmemory load stops new work
Three agent sessions once exhausted the host's memory. Now, under memory load nothing new is dispatched. Past a high-water mark, a governor freezes agent jobs and thaws them one at a time once the pressure has stayed normal. Heavy runs go through a memory-capped launcher.
Sources: the work board, its claim ledger and claim events, read on 2026-10-01 at 19:58 UTC; the dispatcher's audit log, 22 September to 1 October 2026; the seat roster, dispatcher, relaunch tools, skills and coordinator contract on the handbook's main branch. The claim ledger starts on 23 August 2026. Operator leases are the operator's own claims on the same board.