The alert said 60 percent. My usage dashboard for 2026-09-10 was showing spend at nearly double the trailing week’s daily average, and the first instinct was to start hunting for a runaway agent, a stuck loop, a session that forgot to stop. That instinct was wrong, and it cost about twenty minutes before the actual bug surfaced: the number itself was broken, not the spending it was supposedly measuring.
The wrong turn: chasing a runaway that wasn’t there
The work log for that stretch was full of legitimate candidates. A session log entry from that window reads: “Found the runaway: an orphaned issue-tracker rewrite sweep from a closed session was burning credits and popping console windows.” That’s a real incident, and it primed me to assume the quota spike had the same shape. So I went looking for its twin: grepped the wrapper-cost logs for any process still tagged active past its session’s end, checked ops-cli background for zombie PIDs, cross-referenced scheduled task registrations against the headless-migration tracking list to see if something had started re-spawning windows and credits with it.
Nothing. Every session in the window closed cleanly. No orphaned sweep, no duplicate cron, no process holding a lock past its lifetime. I’d spent real time verifying that the thing that actually happened three days earlier (that issue-tracker sweep) wasn’t happening again, when the report I was reading was never describing spend accurately in the first place.
What the rollup was actually doing
The rollup script read each session’s transcript and summed usage blocks per file. Two bugs, both in how the read happened rather than in anything an agent did:
Double-counting. Usage blocks were being read per file scan, and a long-running session’s transcript gets rewritten incrementally as it goes. Without a stable identity per usage event, a block that had already been counted on a prior scan got counted again on the next one if the file had grown. There was no dedup key at all: it just summed whatever blocks were present in the file each time it ran.
Misattribution by mtime. The rollup bucketed cost by file modification time, not by the timestamp inside the message. A session that started at 11 PM and kept writing to its transcript past midnight had its early spend filed under the day the file was last touched, not the day the tokens were actually spent. Any session crossing midnight, and with sessions running the kind of hours in that day’s log (18h12m, 16h17m, 15h14m durations, several starting mid-afternoon and running to 5 or 6 AM), smeared its cost onto the wrong day almost as a rule rather than an exception.
Put together: a session’s early spend got filed on day N, some of its later blocks got recounted on a subsequent scan and filed on day N+1 by mtime, and day N+1 inherited both its own real spend and a phantom slice of the day before. That’s how an ordinary day reads as a 60 percent spike without a single extra token being spent.
The fix
We rewrote the rollup around a SQLite-backed store instead of a re-scan-and-sum:
- Usage blocks are deduped by
message.id, so a block that already has a row never gets summed twice regardless of how many times the source file gets rescanned. - Buckets are keyed by the message’s own timestamp, not file mtime.
- The day boundary moved to 6am to 6am instead of midnight to midnight, because a working day in this shop routinely starts in the afternoon and ends past 3 or 4 AM. Midnight was never a real boundary here; it was just the boundary the naive timestamp math picked by default.
- Each file gets an incremental per-file cursor, so a rescan picks up only new blocks instead of re-reading and re-deriving the whole file.
- The wrapper-cost log entries fold into the same rows as the primary transcript usage, so a session’s total isn’t split across two places that could drift independently.
Shipped as ops-cli usage-rollup {backfill,update,status,days,sessions,messages}, backfilled 89 days back to 2026-06-13, with a new test suite (253 lines) pinning the dedup and bucketing behavior directly so a future rewrite can’t reintroduce the same double-count silently.
Same week, same lesson, different layer
The quota bug was a measurement problem in software. The same week, a separate investigation turned up a measurement problem in hardware, and it’s worth naming because it’s the identical failure shape one layer down. The machine had been running slow, and the assumption going in was the usual one: some single process is pegging the CPU, find it, kill it. Instead: “Found the machine was creation-bound rather than compute-bound, with no single hot process: a dozen ordinary things each spawning handfuls of processes across thirteen live sessions, while every diagnostic tool failed silently under exactly the load it was meant to measure.” The process queue was sitting around 300. After fixing the safety guards that were failing open and collapsing five processes per edit down to one, the queue dropped to near zero, no single culprit ever identified because there wasn’t one.
Two incidents, one week, same root shape: the thing reporting on the system was failing precisely under the load the system was actually generating, and in both cases that failure looked, at first glance, exactly like the problem it was supposed to be catching. A quota tool that double-counts under real session volume and a diagnostic tool that goes silent under real process load are the same bug wearing different clothes. Once you’re running agents at volume instead of running one script at a time, your instrumentation stops being a spectator. It’s failing under the same conditions as everything else, and it has to be debugged the same way.