---
title: "The Small Annoyance That Exposed Our Capacity Ceiling"
canonical: https://dxdev.com/blog/2026-09-18_popup-popups-reveal-architecture-ceiling/
datePublished: 2026-09-18
---
A powershell.exe window opened on top of my editor every time any agent session finished a turn, and on the morning I chased it I had six sessions open at once.

## One line, one conhost per turn

The Stop hook in `settings.json` runs on every turn end, across every concurrent session. It played a wav file, and it did it inline through `"shell": "powershell"`:

```powershell
(New-Object System.Media.SoundPlayer 'C:\Windows\Media\Windows Ding.wav').PlaySync()
```

When a console-app child has no console to inherit, Windows allocates it a fresh one. I sampled the process table live while a turn ended. The spawned powershell.exe had its own conhost.exe, and there was no hiding flag in its argv. The ding was fine. The window was a side effect of how the hook launched.

The fix was to move the line into its own `.ps1` file and launch it the way every scheduled task in the repo already launched: `wscript.exe //B run_hidden.vbs <script>`. Two other spawn paths turned out to be missing the same protection. For Python callers we already had a helper, `no_window_flags()`, which returns `{"creationflags": subprocess.CREATE_NO_WINDOW}` on win32 and `{}` elsewhere. Both got the flag.

## I had already fixed this bug twice, one script at a time

This was not a new failure, and the earlier history is the embarrassing part.

A cron task that sweeps our chat channels every five minutes was first registered as a direct powershell.exe action. It flashed a console on top of my work every five minutes for two days. `-WindowStyle Hidden` does not help with this class of problem: conhost creates the window before PowerShell gets a chance to hide it. So I moved that one task onto the wscript wrapper and wrote a comment in the script explaining why.

That was a fix for one script. On September 11 the same symptom came back from a `pythonw` sweep script spawning `claude -p` children with no creation flag, and I fixed that one script too, this time with the helper. I also wrote an audit that flags any interactive scheduled task launching powershell, cmd or python without the wrapper. It reads scheduled tasks, and a Stop hook is not a scheduled task, so it never looked at the hook that fires the most often of anything I run.

Seven weeks passed between the first fix and the hook. In that time I paid for the same mistake three times, because each fix was scoped to the file where I happened to see the window.

## The popup rate was a load meter

Once I stopped treating it as an annoyance, the frequency was the interesting part. The hook fires once per turn end per session, so the popup rate was the sum of turn rates across everything running. The work log for that morning shows six sessions open together at 8:30, started at 12:04, 1:14, 2:09, 2:13, 8:12 and 8:27, the oldest with more than eight hours on it. I had built a lot on top of the assumption that one machine and one Claude account could carry that. Nothing had ever measured whether they could.

There was still no counter for the machine. When sessions slowed down together the next day I had nothing to look at. So I added a load sampler that writes one line every 5 seconds. Since it went in, CPU has been above 85% in roughly 39% of samples while RAM sat near 52%. My working hypothesis is CPU saturation from concurrent agent work and process churn. It is unproven, and the log says so.

The account side had a worse blind spot. Our unattended daemon spent hours retrying a Claude seat that had run out of weekly quota. A usage limit arrives on standard output with nothing on standard error, so my failure classifier, which watched stderr, never registered it as a quota failure. It just looked like a slow call worth retrying. Three fixes followed:

- The router treats a seat at 95% weekly usage or higher as dry.
- It picks a seat with quota left before each call.
- The pacing sampler now fails loudly when a seat goes invisible to routing. It quietly reported success before, and it actually hit that case the same morning.

That week also included a hard ceiling event that forced one project's work onto a different account entirely.

## The business case for a second account

Once the ceiling was a logged event and no longer a vibe, the request to a teammate to fund a second account was easy to write. I put the plan cost next to what the agent work had returned and got roughly 70x. The ceiling event went in as the evidence that we were already paying the cost of being capped, just in a form no invoice shows.

The burn rate itself is still the open problem. A second seat raises the ceiling and does nothing about the rate at which six sessions climb toward it.

## Admission control before the spawn

The gate I built on the sampler is an admission check, and it runs before the handoff-spawn command starts anything, before the seed is stored or the seat is pinned, so a refusal leaves nothing half-started. It makes two checks. One is how many spawned sessions are RUNNING, using the count the subsessions command already had. The other is the last window of sampler data: mean CPU over a threshold, RAM under a floor, or commit charge over a ceiling.

Two rules came out of the failures above. If the sampler log is missing or stale, the check admits and says "no pressure data", because refusing when the instrument is off turns an observability gap into an outage. And hub and lane-manager spawns are never refused, only reported. Executors and task managers get deferred, and each deferral is appended to a `decisions.jsonl` log with the numbers that caused it.

The first capacity meter I owned was a sound effect with a side effect. It worked because it was annoying enough that I finally went looking for the cause.
