The queue depth read 300. That was the number that made the slowness real instead of anecdotal, and it was also the number that turned out to be a symptom of something completely different from what I went looking for.

Thirteen live sessions were running when the machine started crawling: my primary coding session, its mirrored agent counterpart, a batch judge sweep, a grounding pass, a hub-prompt session, an item review. My first assumption was compute. Something in that stack was pegging a core, probably one of the haiku sessions running long diffs (one of the judge-sweep tasks alone burned 294 seconds on 23,602 tokens). So I went hunting for the hot process. top showed nothing pinned. No single process over a few percent CPU, sustained. That ruled out the story I’d walked in with.

So I reached for the diagnostic tools built for exactly this: process-count-over-time sampling, a spawn tracer that logs every child process a task creates. Both came back clean, or close to it, under load that was visibly making the machine unusable. That was the wrong turn, and it cost real time. I trusted the tooling because it was purpose-built for this, and it took most of a session cycle before I noticed the tools themselves were failing silently under the exact load they were supposed to measure. The spawn tracer was itself spawning a subprocess per sample. Under thirteen concurrent sessions, its own sampling overhead was part of the noise it was trying to isolate, and it never surfaced that fact. It just returned thin, unremarkable data and let me conclude nothing’s wrong here.

Once I stopped trusting the sampler’s silence and started counting processes by hand, per session, the shape flipped. It wasn’t one hot process. It was a dozen ordinary things, each spawning a handful of child processes, multiplied across thirteen concurrently running sessions. Guard hooks fired per edit, and each one was shelling out separately instead of batching: five processes per edit event, not one. With sessions editing files continuously, that’s five spawns times every edit times thirteen sessions running in parallel. None of them individually looked wrong. The aggregate was the outage.

The deeper issue sat inside the safety guards themselves. They’re supposed to fail closed: block the risky action, hold the process, require confirmation. Under contention they were failing open instead. When a guard hook couldn’t get a lock or timed out waiting on a resource another session held, it let the action through rather than blocking it. That’s the design flaw that turned a bit slow into creation-bound with no single hot process to point at. A guard that fails open under load doesn’t just stop protecting anything, it becomes invisible while doing it, because the failure mode looks identical to success from the caller’s side. No error, no log line, just a spawned process that shouldn’t have spawned.

The fix was threefold. First, collapse the five-processes-per-edit guard hook path into one process instead of shelling out separately for each check. Second, make the fail-open behavior explicit and reversible: a deliberate off switch for guards that actually persists, including across restarts and maintenance windows, instead of silently reverting to permissive the moment something else went wrong. Third, and this is the part that mattered for next time, extend the audit itself. Two new categories: a process with no owning session left alive, and a process running with no discernible purpose tied to any active task. Before this, an abandoned or purposeless process could sit there burning a core indefinitely and nothing would flag it, because the audit only knew how to ask “is this process doing its declared job,” not “should this process exist at all.”

After the fix landed, the processor queue went from roughly 300 back down to near zero. That’s the number that confirmed it, not the CPU graph, not the spawn tracer, the actual OS scheduler queue depth, measured before and after.

Both the guard hooks and the spawn tracer had the same blind spot: neither accounted for what happens to itself when the system it’s protecting is under the exact load it exists to catch. A guard hook and a diagnostic sampler both looked healthy in isolation and both were lying, for the same underlying reason. The audit has categories for orphaned and purposeless now because not obviously broken turned out to be a category that included the actual outage.