---
title: "We Built Automatic Quota Failover, Then Ran a Drill and Watched It Miss Five Crons"
canonical: https://dxdev.com/blog/2026-09-12_failover-tested-incomplete/
datePublished: 2026-09-12
---
Over 60 percent of a week's Claude quota vanished in a single day. I went looking for a runaway loop and didn't find one. I found two ordinary causes stacked together. One was a wide spread of daytime sessions. The other was a large unattended overnight run I hadn't been watching.

That morning I built a usage-and-outcome tracking page, so the next spike shows up as a chart instead of a guess. But a tracking page only tells you what already happened. The bigger problem underneath it was that the whole setup ran on one account. One account is one point of failure. If that account runs dry mid-afternoon, every session and every scheduled job behind it just stops, silently, until someone notices a stale ticket.

## A second account, and checking it actually worked

The businesses we run share this setup, so the obvious move was a second account. I did not trust that the rest of the tooling would carry over untouched. I asked the new login whether it still had the same skills and access, then diffed the two account configs side by side.

The diff found a real gap. The new login was missing the browser-control tool and project trust, both silently absent rather than erroring out. Trust settings and tool grants turned out to live per-account, not per-launcher. I copied the config over by hand and verified a fresh process could reach all three connected tools before trusting it with anything real. That was a separate side task, about 22 minutes, not time out of the failover work itself.

## Making the account switch invisible to the scheduled jobs

With two working accounts, the actual goal was for sessions and scheduled jobs to report which account they're spending, so the usage guards could watch both, and for alerting to route by business instead of dumping everything into one channel. I split it into four per-business sweeps and wired in automatic failover: when an account's quota guard trips, the jobs on it should hop to the other account without anyone touching a config file.

That's the part I didn't verify by inspection. I ran it as a drill.

## What the drill actually found

The test was simple: force one account into a dry state during the pre-close window, when the job load is heaviest, and watch which ones picked up the failover and which didn't. Nine of fourteen jobs paged over cleanly. The other five didn't move.

The five that stayed dark never reached the failover path at all. Four of them were alert jobs that exited with a success code no matter what actually happened inside. The check that decides "this account is dry, fail over" never ran, because nothing ever triggered it. A fifth job had the same shape. A separate script sitting next to all of them was reading a cached copy of which account was currently active rather than checking fresh, so even a clean failover elsewhere could still look stale to it.

That gap is now an open item on the task itself, not a footnote in a retro. The instinct after a drill like that is to say "close enough, most of it worked" and move on, but nine out of fourteen isn't a passing grade when the failure mode is a business going dark during a busy week with no page at all. The fix is mechanical: make the four alert jobs and the fifth actually surface their real exit code instead of swallowing it, fix the script reading the stale account state, and re-run the same drill against all fourteen before calling it done.

Infrastructure built to catch failures needs its own drill before you believe it. Run it in production, under the load it was built for, and check every job it is supposed to cover instead of assuming the newest ones represent all of them. The tracking page told me what already burned. The drill told me what would have burned next.
