One AI session burned through 365,257 output tokens and 37,945,352 cache reads, which are bits of earlier text pulled back into the conversation, and none of my rules for avoiding that had a chance to fire. Those rules were to search inside files before reading them, clear old conversations between tasks, and send side searches elsewhere. They all assume I know what I am looking for, and during an audit of session notes and instruction files I did not. I kept reading, comparing, and trying another angle, and the count kept climbing.
A token is a small piece of text that an AI counts. This one session had also pulled in 37,945,352 cache reads, which are bits of earlier text brought back into the conversation. It used a meaningful slice of the quota I had for the week.
That was uncomfortable because I already had rules for avoiding this. I had told myself to search inside files before reading them all. I had told myself to clear out old conversations between tasks. I had told myself to send side searches elsewhere so the main conversation would stay tidy.
Those were good rules. They worked when I knew what I needed.
The afternoon that broke them was an audit. I was reading session notes and instruction files, comparing results, and revising a small script. I did not know the problem at the beginning. I had to look around before I could even name the right question.
That is where my plan fell apart. You cannot search first when you have not yet learned what needs finding. I kept reading, comparing, and trying another angle. The count kept climbing. The constraint was real, but the strategy was wrong for that kind of work.
So I tested RTK. It is a proxy, a small go-between that looks at what ordinary computer tools send to the AI before the AI has to read it. A clean status check can become a short summary. A test run can hide the long list of passing tests and show the failures. A file comparison can lose surrounding lines that rarely affect the decision.
I was not looking for a trick that would make every task cheaper. I was looking for a guardrail that would still work when I was tired, moving fast, or half lost in a problem.
The most reassuring part was what happened when the tool did not recognize a command. It left the result alone. No broken command. No mangled answer. Just the full output. That matters because the danger of a new habit is often the moment it gets in the way. Here, the worst case was no savings.
I added one simple instruction to use the filter before routine computer tools. It is still a habit, and habits can be missed. But the tool itself does the shortening once it is used. It does not depend on me remembering a list of careful behaviors in the exact moment those behaviors are hardest to follow.
I have not proved what this saves across a normal week, and the tidy percentages came from test runs and other routine output, where the filter has the most noise to strip. The audit session is the opposite case: it was mostly me reading large files and comparing results, so the honest guess is that it saves less there. Even so, that session put 365,257 output tokens and 37,945,352 cache reads on the meter before I could name my question, and no rule of mine could have searched first for something I had not found yet. The filter does not need me to know what I am looking for. It shortens the status checks and passing-test lists on the way in, and when it does not recognize a command it hands back the full output, so its worst case is the same bill I already paid.