My AI wasn’t ignoring me. Its memory file was lying to both of us.

For weeks I had the same quiet frustration. I’d tell my agent to remember a goal, an idea, a preference. It would say it saved it. And then the very next session it would act like the conversation never happened. I’d repeat myself, it would save again, and the cycle continued. I started to assume this was just what long-running AI was: confident in the moment, amnesiac the next morning.

I was wrong about the cause, and the real one is worth a post because almost everyone building on top of these tools will hit some version of it.

What was actually happening

My setup gives the agent a persistent memory: a plain text file it reads at the start of every session and appends to as it learns things about me and my work. Simple, durable, version-controlled. The kind of thing that feels obviously correct.

The problem is that the harness only loads the first 200 lines (or 25KB, whichever comes first) of that file into context at startup. Everything past that line never reaches the model. No error. No warning I noticed. The file on disk was complete and correct. The agent just never saw the bottom of it.

And because new memories were appended to the end, the newest things I saved were exactly the ones falling off the invisible edge. The agent wasn’t ignoring my goals. It literally could not see them. From its side, they had never been written.

This is the worst class of bug: not a crash, but a silent gap between what you believe is true and what the system actually does. The text was there. The promise that it would be read was the lie.

The fix: two tiers and a guard that yells

Once the cause was clear the fix was almost obvious, and it generalizes.

Split the store into two tiers. A curated, always-loaded file that holds only the core rules and high-frequency facts, deliberately kept under the load budget with headroom. Then a complete, read-on-demand catalog that holds everything (mine is up to 257 entries) and is never auto-loaded, so its size doesn’t matter. The agent pulls from it when a topic comes up. Always-in-context and fetch-when-relevant are two different jobs, and cramming them into one file is what broke.

Make the silent failure loud. I wrote a small checker that fails with a non-zero exit the moment the curated tier creeps over budget, or the moment a memory exists that nothing in the catalog points to (an “orphan” you can only find by luck). The point isn’t the script. The point is converting an invisible degradation into an error that stops me. If a system can silently drop your data, the first thing to build is the thing that refuses to let it.

The lesson worth keeping

Every byte you hand an LLM competes for a fixed budget, and the boundary is usually enforced quietly. Truncation, summarization, eviction: most of it happens without a stack trace. So two rules I’m taking forward.

First, never trust “it’s saved” to mean “it’s loaded.” Those are different claims, and the gap between them is where weeks of frustration hide.

Second, separate what must always be present from what can be fetched on demand, and budget the first tier like the scarce resource it is. The agent doesn’t need everything in front of it at all times. It needs the few things that are always true, plus a reliable way to reach for the rest.

I’d been blaming the model for being forgetful. The model was fine. I’d just built it a memory it was structurally unable to read past line 200.