1,016 open tasks went in and 201 came out closed. It took under two minutes of logged session time, and I don’t trust any of those 201 closes yet.

What the sweep did

The hub session pulled every open task from two queues and triaged each one. Anything that looked finished got a second pass before it was proposed for closing. That second pass is where the accuracy comes from, because “looks finished” is the guess most likely to be wrong. A ticket whose branch merged is finished. A ticket whose title sounds like a shipped fix might be finished, or might be one of six hotfixes that still hasn’t been verified on master.

I approved the 201 closes. Each one was written into a close log at the moment of closing. An entry holds the ticket, the evidence the agent cited for closing it, and the approval. The log exists so that later we can grade the agent’s triage against what actually happened. If a closed ticket gets reopened, or a duplicate is filed within a few weeks, that is a miss, and the log tells us which kind of evidence produced it.

Closing tickets is a one-shot job, and I had no way to know afterward whether the agent was good at it. The log is what lets me find out.

Why the log has to be written at close time

The same day showed me what reconstruction costs. At 8:25 AM the editor hosting our sessions crashed. Five session records were closed without a proper close, and between them they left 43 work units marked open. Later, a cockpit restart at 5:42 PM killed more sessions mid-work. One died after editing two test files and starting a baseline run. Another died with a finished feature uncommitted. The hub had to go back through manifests and finish or hand off each one.

Reconstructing what a dead session meant to do is slow and partly guesswork. A close log written in the same step as the close has none of that problem. It is boring, and it survives crashes because it is not held in anyone’s context.

A wrong turn from the same day

While the hub was running executor sessions as children over the bridge, I had it wait on their replies with /loop and ScheduleWakeup. It seemed the obvious way to poll. It swallowed the replies in the editor, so the hub kept waiting on messages that had already arrived. I wrote the lesson into memory and moved to direct hub-to-child messaging over the bridge, which I then tested in both directions before trusting it.

It is the same failure the close log guards against. An agent believed something had happened, or not happened, and nothing recorded the truth to check against.

A smaller instance came at 2:51 PM. I fixed a red CI run by committing the missing half of a session eviction notice. The commit message cites the wrong ticket, copied from a comment in the work-in-progress code. Nothing broke, but a close log built from commit messages would have attributed that fix to the wrong task. So the log records the evidence the agent used and the ticket it thought it was closing, and grading compares the two.

What the grades will feed

Triage accuracy only matters if it changes something. The workflow fixes I proposed on a follow-up ticket, so that tasks arrive agent-ready, are not built yet. I want the grades to decide what gets built. If the misses cluster around tasks with no acceptance criteria, that fix gets priority. If they cluster around tasks whose evidence lives in another repo, the fix is different.

Same day, a pre-release scan gained a check that flags tickets which were hotfixed but never closed. That is the mirror image of the sweep: one tool finds tasks that should be closed, the other finds closes that should not have happened. Once both have a record, I can compare them.

I have a number for how many tickets I closed and none yet for how many I should have. The log will produce that second number, and the next sweep will be designed around it.