Forty-three windows, closed by hand
Forty-three: that was the number that kicked off the post-mortem. One day, I’d closed forty-three agent windows manually, one at a time, no batch action, no shortcut. When I asked a session to reconstruct what had actually happened that day, it counted something else first: about half of everything I’d typed into any session that day was some version of “where are you.” Status checks. Not instructions, not corrections. Just me asking sessions where they stood.
The first read on that pattern is trust. I don’t trust the agents to tell me when something’s done, so I check constantly, and the checking itself generates the mess I then have to clean up by hand. It’s a plausible story. It’s also the wrong layer to fix, and the post-mortem only found that out by refusing to stop at the plausible story and going to pull the actual instrumentation instead.
What “delivered” actually meant
The mechanism under those forty-three windows was a delivery step that placed a window onto my screen and logged it as done. Placed, not seen. If I wasn’t at the keyboard when a session finished, the window still went up, still counted as a successful handoff, and then sat there in an empty room until I came back, noticed it, and closed it. Every one of those was a false positive: the metric said “delivered,” the ground truth was “nobody was there.”
That’s the same shape as a monitor that looks clean because it’s quietly excluding the thing that would make it dirty. The fix wasn’t “make the agents better at knowing they’re done.” It was presence detection: sessions now check whether I’m actually at the keyboard, and if I’m not, they hand back a link instead of placing a window nobody will see. The “success” event only fires when there’s someone there to receive it.
The watcher that had been lying for two days
Digging into the attention data to size the problem turned up a second, uglier finding. The watcher responsible for tracking whether I was present at the keyboard had been running stale code for two days. Not down, not erroring, just silently executing an old build the whole time. Which meant any conclusion drawn from “what attention data showed for the last two days” needed to be thrown out before it could be used as a baseline for anything, including the post-mortem that was currently trying to use it.
Same day, same audit, a third instrument turned out to be estimating instead of measuring: the trust ladder, the thing meant to track how often I approved a session’s output, was guessing at roughly 21 percent of my approvals rather than counting them. A number that looked like ground truth was partly invented.
None of these three were the same bug. But they were the same failure mode: a system reporting a clean signal that wasn’t actually derived from the event it claimed to represent. Fix the top-line symptom (I close a lot of windows by hand, so add more automation to close windows) and you’d have shipped something that made the dashboard look better without touching why the windows were empty in the first place.
The fixes that actually shipped were boring by comparison: presence-aware delivery instead of blind window placement, the desk queue stopped hoarding requests from sessions that had already closed (another source of phantom backlog inflating the “still pending” count), the watcher got put back on live code, and the trust ladder started counting real approvals instead of estimating them. A full handoff went onto the ticket for whoever picked up the thread next.
The pattern showed up again the same day
A separate ticket filed that same session was for an order-dependent test hiding inside the suite’s own baseline: a test whose pass or fail outcome depended on what ran before it, quietly living inside what the team treated as “green.” Different subsystem, identical shape. The baseline you’re diffing your fix against has to be verified as a baseline before the diff means anything. A metric that improved because the thing measuring it changed underneath you isn’t a fix. It’s noise wearing a fix’s clothes, and the only way to tell the difference is to go open the instrument itself before you believe what it’s reporting.