The weekly total came back as 46.9M output tokens and 9.8 billion cache reads, and I read the two numbers twice because I assumed I had misplaced a decimal. I hadn’t. Over seven days, for every token an agent wrote, about 209 tokens were read back from cache.
Every agent session I run leaves a JSONL transcript, and each assistant turn carries a usage block with four counters: input_tokens, output_tokens, cache_creation_input_tokens and cache_read_input_tokens. I had never summed them. The daily work log already prints them per session, so I wrote a small script to walk the transcript folders and add up the same four fields for the week.
The wrong turn: I shrank the model
Before I looked, I had a theory. The bill felt too high, and the obvious lever was the model. So I acted on it. The scheduled jobs, the ones that fire claude -p every half hour to classify backlog items and produce the morning digest, all went down to Sonnet at medium effort. My interactive work stayed on the bigger model, but I started tolerating weaker output there too, because cheaper felt like progress.
That cost me two things. First, quality. The classification workers started disagreeing with me, and one afternoon I moved five tickets between boards on the strength of a classification rule that was wrong, then had to reverse three of them by hand. The rule was my error, not the model’s, but I had gone looking for savings in exactly the place where a small mistake gets expensive. Second, the change did nothing to the number I was worried about. When I finally did the arithmetic, the scheduled workers were the cheapest thing on the machine.
Look at one of those workers in the log. A single classification run: 1 prompt, about 20,000 output tokens, 51,766 cache-read tokens. Run that every 30 minutes and you get roughly 2.5M cache reads a day. That is a rounding error.
What was actually big
Here is one interactive session from the same day, straight from the log:
1 session, 397 promptsin=1,376 out=425,998 cache_read=168,910,001And the session that made me write this post: 483 turns, 128M cached tokens read against 428k of fresh input plus output combined. That is a ratio of roughly 300 to 1 inside a single conversation.
The mechanism is not subtle once you say it out loud. Each turn re-sends the whole conversation so far. The provider serves most of that prefix from cache, which is why it is cheap per token, but “cheap per token” times “the entire history” times “every turn” is the whole bill. The cost of turn N is roughly proportional to everything said in turns 1 through N-1. Sum that over 483 turns and you get an area under a curve that grows with the square of the thread length. The fresh tokens I typed and the tokens the model wrote are a thin sliver on top.
Averaged out, that 483-turn session was reading about 265k tokens of context on every single turn. Most of it was old tool output: file dumps, grep results, log tails from things I had already dealt with two hours earlier and never needed again.
Why the model swap couldn’t have worked
Cache reads are billed at a fraction of the normal input rate, and a cheaper model discounts that rate further. Fine. But a discount on a multiplier doesn’t change the multiplier. If a thread is 10 times longer than it needs to be, you pay for that on any model. Moving my scheduled jobs down a tier shaved a little off a line item that was already tiny, while the long interactive threads kept accruing quadratically.
The lever that matters is thread length times turn count. Halving the model price halves the bill. Cutting a 483-turn thread into four threads of roughly 120 turns each does far better than that, because the sum of average context sizes drops with the square, not linearly. Four short threads each carry a small history. One long thread drags its whole history through every turn.
What I changed
Nothing clever.
- Close threads at a natural seam. Ticket shipped and verified, thread closes. The work log entry becomes the handoff, not the transcript. This day’s log has 106 sessions in it, and the ones that finished cleanly are the cheap ones.
- Write state down, then start fresh. When a task has to survive a break, I write the resume state into the ticket (what is done, what is parked, what decision is pending) and open a new session that reads only that. A page of notes replaces a quarter-million tokens of scrollback.
- Stop dumping raw output into the thread. A tail of a log I need for one decision gets read once and summarized in a sentence. It doesn’t have to ride along for the next 300 turns.
- Put the model back. The interactive work is on the model I trust for it again. The scheduled workers stay on Sonnet because they are cheap and their job is narrow, which is a fine reason and not a savings plan.
The number I watch now
I added the aggregation to the daily sync so every session line already shows in / out / cache_read. The column I look at is cache_read divided by turns. When one session’s figure climbs past a couple hundred thousand per turn, that thread has outlived its usefulness, and the right move is to write down where I am and close it.
The 9.8 billion figure doesn’t scare me anymore. It is mostly a measure of how long I let conversations run, and that is a habit I can change without touching a single model setting.