Thirteen cards, five real
The secretary dashboard was showing thirteen active-work cards. Five of them were real sessions. The other eight were zombie manifests, leftover session files that never got cleaned up when a run ended, still rendering as if an agent were mid-task. We wrote a filter to cut the noise: 13 cards down to 5, and closed a related duplicate-manifest bug where the same session was spawning two ghost entries on the board.
That was the surface bug. The real problem was underneath it, and the filter didn’t fix that part.
Eight of thirteen cards, nearly two-thirds of the board, were noise I had been treating as real. I do not know how much time I’ve spent clicking into a stale manifest expecting a live session and finding nothing, but it happened often enough that a filter felt worth writing. I trusted that board’s status column for longer than I should have before asking whether the thing setting the color could be wrong in the first place.
What “amber” actually meant
Cards on the dashboard have a status color. Amber means: agent says this needs your attention, come look. The dashboard was rebuilt that day into something closer to a real operating surface, live per-agent progress inside each fleet card, drill-in on demand, spawn-from-task so a worker can be kicked off directly from a ticket instead of a terminal. All of that made the board more useful to look at. None of it changed what amber was actually reporting.
Amber was self-report. A worker finished a task, decided for itself that the task was done, and flipped its own card to amber so I’d come review it. There was no step between “the agent thinks it’s done” and “it’s now sitting in my queue demanding a decision.” Every card asking for judgment was, structurally, just an agent’s opinion of its own output. On a day with 33 sessions and 107 commits across 14 repos, that opinion is the bottleneck. If half of what lands in amber turns out wrong on inspection, I’m not reviewing work, I’m doing the verification the worker skipped, one card at a time.
I had been running this dashboard for weeks before I asked that question directly. Amber meant “come look” and I came and looked, on the assumption that a worker declaring itself done was at least a reasonable signal for where my attention belonged. I never had a real count of how often that signal was wrong until the zombie-manifest cleanup forced me to look at every card on the board at once, and finding eight fakes out of thirteen is what made me stop trusting the five that were left.
The contract change
We’d already been doing review cascades manually elsewhere. The docs work on the main product that day went through one model’s review pass, then a second, independent model’s pass, before it shipped. The blog redesign went through the same pattern, one review round, then a second, independent round, before going live. Those worked, but they were bolted on after the fact, a separate step someone had to remember to run. The dashboard’s amber problem was the same failure in a different shape: verification treated as optional, added back in only when someone thought to add it.
The fix came from a paper on agent-loop design, out of one of the frontier-model labs, that argues a spawned worker should carry two things from the start, not one: a rubric defining what success looks like, specific enough to check against, and a verifier, a second pass independent of the worker’s own claim, that confirms the rubric was actually met. Not “the agent says it’s done.” “The agent’s output was checked against a stated bar, by something other than the agent asserting it.”
We rebuilt the spawn seed, the initial task definition a worker gets handed, to require both fields. No rubric, no spawn. And we added a footer hook that runs at session close: verify-before-amber. A session cannot flip its own card to amber. It can only request a verification pass, and the card only goes amber once that pass returns a result, not once the worker declares victory.
Why not just review everything after the fact
The obvious alternative was to keep doing what we were doing: let workers report done, and run review cascades manually on the important stuff. We rejected it for the same reason the zombie manifest bug existed in the first place. Manual steps get skipped when volume goes up, not when it’s convenient. Eight stale manifests sat on the board because nobody had a mechanism forcing cleanup, they just accumulated. An after-the-fact review step has the identical failure mode: it works until the day is busy enough that someone forgets to run it, and busy is exactly when you need it most.
The other alternative was to make the worker verify itself, harder, more test cases, a self-check pass before reporting done. We didn’t trust that either, and for the same underlying reason a zombie manifest doesn’t self-report as stale: a process checking its own state has no independent signal when that state is wrong. The verifier has to be a separate pass with its own read of the rubric, not the same agent grading its own homework twice.
What changed on the board
Sessions endpoint response time dropped from 42 seconds to 11 milliseconds in the same rebuild, a TTL cache plus single-flight fix that was overdue regardless of the rubric work, but it mattered here because a faster board is what makes verify-before-amber tolerable. If checking a worker’s status is slow, adding a mandatory verification hop on top of it would have made the dashboard worse to use, not better. Fixing both at once meant the extra step didn’t cost visible latency.
The doctrine itself is written up as a standing note in the vault (2026-06-11_loop-design-applied) rather than left as tribal knowledge in one commit message, so the next spawn seed anyone writes inherits the same contract by default instead of by memory.
The eight zombie cards are gone. The five real ones are still there. The difference now is that when one of them goes amber, it’s because something checked it, not because it checked itself.