---
title: "I Built a Usage Counter and It Inflated Itself by Reading Its Own Code"
canonical: https://dxdev.com/blog/2026-09-26_the-metric-that-read-its-own-source-code/
datePublished: 2026-07-30
---
The counter said my subagents had burned exactly three times the tokens they had. Exactly three, to the digit, which is what a real measurement almost never gives you.

## The wrong answer I gave first

The question was simple. The token report for a session covers the main transcript, so are the subagents it spawns counted anywhere? I checked, didn't see them, and answered no. That was the first mistake, because I never looked at where subagents write their output.

Instead I built a scraper. The main transcript contains a completion notice whenever a subagent finishes, wrapped in a tag and carrying a usage block. So the plan was to walk the transcript, match the tag, pull the token figures out of each block, and add them to the session total. It took an afternoon and it ran clean on the first try. Clean is what worried me later.

I pointed it at a session where I knew roughly what the delegated work should cost, and the subagent line came out at three times the plausible figure. I did not doubt it at first. I assumed I had underestimated the fan-out, and I wrote the number into the report.

## What the three was

I only started diagnosing when I diffed the scraper's matches against the subagent count I could see with my own eyes. Three subagents had finished. The scraper reported nine notices.

The harness does not write a completion notice once. The same notice shows up on several lines of the transcript: once as the queued event, once as the message injected back into the conversation, once again in the record of that turn. Same tag, same usage block, three places. My matcher treated every occurrence as a separate subagent completion. Nothing was randomly wrong. Every subagent was counted three times.

The obvious repair was to dedupe on the notice's task id. That cut the inflation, but the total was still above what I could account for by hand. That was the second wrong turn, and it cost me another evening of patching a number I had already published once.

## The tag that matches itself

The leftover excess came from somewhere stranger. The tag I was matching is a plain string, and plain strings appear in prose. Every time a conversation mentioned the notice format, the transcript contained that tag as ordinary text. That included the sessions where I was discussing the scraper, and it included the source of the scraper module, which contains the tag as a literal in its own regex, and which I had read into context several times while debugging.

So the tool was billing itself. Every conversation about how the counter worked created new matches for the counter to count. The more I worked on it, the higher the number went. A metric that grows when you look at it is not measuring the thing you pointed it at.

The fix was to anchor on structure: only count the notice when it appears in the specific record type the harness uses for it, at the start of the message rather than anywhere in the text. The excess disappeared and the count matched my hand tally of three subagents.

## The measurement that was already on disk

With the scraper finally agreeing with reality, I went to compare it against something independent, and found the independent source was already sitting on disk. Subagents write their own transcripts. Each one gets its own file with its own usage records, in the same format as any other session. The existing counter walks every transcript under the root, so it had been picking those files up all along. My original answer of "no, they are not counted" was simply wrong.

The scraper had never been filling a gap. It was a second, worse measurement of a quantity the first tool already had exactly, layered on top of it, which is why it double counted so cleanly.

I deleted it. Roughly a day of work, gone, and the total went back to the honest figure.

## The coverage table

What I kept is the thing I should have asked for at the start: a statement of what each number does and does not include. It lives next to the counter now.

| Path | Status |
|---|---|
| Main session transcripts | Exact. Usage comes from the API's own records. |
| Subagent transcripts | Exact. Separate files, already walked by the counter. |
| `claude -p` worker runs (the scheduled classification and triage jobs) | Exact when they leave a session file. Floor otherwise. |
| `codex exec` runs | Floor. The log line carries a token figure, but it is not reconciled against a transcript. |
| Anything that only shows up as a notice inside another transcript | Not counted, on purpose. It is a pointer to a number, and the number is already counted at its source. |

"Floor" means the true figure is at least this, possibly more. I would rather publish a number labeled as a lower bound than one that looks exact and is wrong, because I have now done both and only one of them cost me a day.

## Two habits that came out of it

Both are cheap.

First, when a new counter disagrees with an old one, I reconcile them on a single session with a known answer before I trust either. A count of three subagents that reports nine is visible in about a minute if you compare against a session small enough to count by hand.

Second, I grep for the string I am matching in the source of the thing doing the matching. If the pattern appears in its own module, the tool has a path to counting itself, and the fix has to be structural, not textual.

The three-times number was wrong in a way that looked like a precise measurement. A suspiciously clean ratio is now the first thing I distrust.
