A Claude seat ran out of weekly quota, and our ticket bot kept calling it for hours. Nothing in its logs read as a quota failure, because as far as the bot could tell, no failure had happened.

What the classifier read

The bot shells out to claude -p for each ticket comment, using a seat, which is a logged-in Claude account with a weekly allowance. When a call fails, the retry logic classifies why, so it can back off on rate limits and rotate on exhaustion. Roughly:

result = run_claude(prompt, seat)
if result.returncode != 0 or result.stderr:
kind = classify(result.stderr) # "quota", "rate_limit", "transient"
else:
kind = "ok"

I had written classify against stderr because that is where errors go. I modeled quota exhaustion as an error, the way an HTTP 429 is an error. It is not one. When a seat hits its weekly limit, the usage-limit message is printed on standard output as if it were the model’s reply, and standard error stays empty. So the call went through the “ok” branch, or through the generic transient bucket when the exit status was nonzero. Either way classify never saw the word quota. The usage-limit text got treated as the reply to the ticket, or as a flaky call worth another attempt on the same seat.

Found by accident

No failing check caught this. A separate session was investigating why the agents were spending a maxed-out account without knowing it, and the retry history showed the same seat being hit again and again for hours with no rotation. The watchdog that was supposed to catch a stuck seat was not running at all. That is a second bug, and it meant the first one had no backstop.

The wrong turn was mine and it was not subtle. I had built the monitor, tested it against the failures I could imagine, and shipped it. Every test I wrote produced stderr, because I wrote the fixtures. The test suite passed the whole time the seat was burning retries, since it only proved the classifier handled the shape of failure I already believed in.

Splitting the streams

The diagnosis took one move once I stopped trusting the classifier’s inputs: run the call by hand against the exhausted seat with stdout and stderr captured separately. Stdout held the limit message. Stderr was empty. Every layer above that had been faithfully processing an empty string.

The fix has two parts:

  • classify now scans both streams, and a usage-limit message on stdout is a quota failure regardless of exit status or stderr.
  • The bot now picks a seat with quota left before each call, instead of finding out afterward.

I verified it live on the running bot, not just in tests. A refused issue-tracker comment also now keeps its detail instead of being binned, so the next odd failure leaves evidence.

The sampler that reported success

The pacing sampler is the other half. It polls each seat’s usage so routing knows who has headroom. Its failure mode was the same mistake in a different costume: when a seat dropped out of what the sampler could see, it reported success on the seats it could see. An invisible seat and a healthy seat produced the same green output.

I changed it to fail loudly. If a seat that routing knows about has no reading, the run errors and says which one. The pacing tests went to 42 passing, and I added cases where a seat is missing from the sampler’s view, which is the case I had never written a fixture for.

That same morning, a seat went invisible to routing for real. The sampler said so, by name, instead of averaging over the gap. I had not staged it. It was the first live proof the check could fail.

What is still open

The burn rate is still the real problem. Picking a seat with quota is a floor. Several unattended jobs draw on a small pool of seats, and one heavy day can take a seat to its ceiling. We already had one hard ceiling event that pushed the platform work onto another account. We filed a multi-seat epic for it, and it is not solved.

What I take from this is narrower than “monitor more.” A check only covers the failure you imagined when you wrote it, and I had imagined an error on stderr. The vendor chose a different exit path, and my fixtures, written from my own assumption, agreed with me. Now a seat with no reading fails the run, and I treat silence from a component as the first thing to page on.