The failover ladder passed its fire drill at 1:23 AM, and the system description I built around it was still wrong on four separate claims.

The drill was three calls with the prompt “ladder”. Claude answered in 4.15 seconds, Codex in 4.23, Manus in 22.49. Every rung of the prepaid failover worked. I had spent the evening building a CLI watch (one window over every live session and subagent), a CLI ask that fails over across Claude, Codex and Manus, and a /brief rebuilt on published artifacts. I then wrote a hundred-word description of how the system runs. Every mechanism in it was real and every sentence checked out against the code.

The review that could only say yes

My own review pass asked one question per sentence: does this mechanism exist, and does it behave as described? It did, every time. The description came back clean. That should have worried me, because a review that only confirms the tool does what the tool does can’t tell you whether you described the right thing.

I sent it out anyway. Manus returned a review in 229.84 seconds at zero credits. Codex, at medium effort, took 120.97 seconds and 40,623 tokens, and its answer was a single change: name the control plane itself, instead of naming two of its mechanisms. I had described the parts and skipped the thing that holds them together.

Four corrections, one mistake

Then I read the description to the person who actually runs this, and they corrected four claims. My notes from that night say all four were the same mistake: I was measuring the tooling’s model of the work instead of how they work. I described failover as the story because failover is the thing I built and drilled. They don’t experience the system as a ladder. They experience it as a set of sessions that are either parked waiting on them or not.

My first response was to patch claim by claim. Fix one, reread, find the next. That cost most of the sequence, because each patch used the same source as the original error. The tool’s own inventory can’t tell you which of its features matter to the person using it, however many times you re-derive it.

What worked was reading their habits over eight weeks from the session logs and rewriting the description around what they do. That read was the external eye I should have started with, and it was aimed at the workflow instead of the tool.

The same blind spot, in daytime

The pattern came back later that same day, in other places.

The plan that was already built. A plan for a live control surface opened with status: PLAN for review. Not built. A multi-review cascade found the opposite. server/stateModel.ts (29 KB) and server/stateStream.ts (the SSE channel) already existed. A 470-line Now.tsx was already routed at /. Probing 127.0.0.1:3010, /api/spine and /api/state returned 403, because the route existed and wanted a key. An unknown path returned 200, the SPA catch-all. A route that didn’t exist would have fallen through to the catch-all. The plan described an intention. It had never been checked against the running system, and Codex reached the same conclusion from a different direction, reading App.tsx and stateStream.ts.

The pool that was sound and unused. At 10:19 AM a Codex review of our Chrome pool came back with this: the pool architecture is sound, but it is not the default execution path. Manus, 630.93 seconds later, said to keep the architecture and add fenced ownership, eligible-only eviction and live-resource validation. Two reviewers agreed the design was fine. The finding that mattered was about which path the work actually takes, and the design doc could never have shown that.

The dev server behind the 502s. In the afternoon, one internal host was being served by a tsx watch dev server. Every file save restarted Node, and anyone loading the page in that window got IIS’s own 502 page. 78 of 12,866 requests that day, and one of them was on a phone. The code was fine and the process supervision was the problem. We switched to node dist/index.js with NODE_ENV=production on :3010, and made an explicit PORT refuse to walk to another port instead of silently 502-ing behind the proxy.

Reviewers see what they are handed

In all four cases a review saw only what it was handed. An internal review gets the artifact and the author’s framing of it, so it verifies the artifact against the framing. An external reviewer with no stake in the framing asks what the thing is for and where the work actually runs, and gets there much faster.

Two rules changed. Any description of the system now gets a pass from a reviewer who has never seen my build notes, and that reviewer is asked which sentence would surprise the person who uses it. And before I review a design, I look at the default path: what actually executes when nobody opts into anything.

A passing fire drill and a passing review both tell you the parts work. Neither tells you the parts are the point.