The memory index was over its size target, so I went looking for duplicates to delete. I found zero.

That was the whole plan. The pile of memory records that every agent session loads had grown past the size we set for it, and the obvious cause was the same fact written down twice by different sessions on different days. Dedupe, shrink, done. I pointed a pass at the pile to find near-identical records.

Nothing came back. No two records said the same thing. I read that as a clean bill of health and nearly closed the task, which would have left the index over its size target with nothing to show for it.

What the empty result was hiding

Grouping the records by subject was the useful part. Once records were clustered by what they were about, I could read each cluster top to bottom, and the clusters weren’t full of copies. They were full of records that disagreed.

Two of those disagreements were worse than a stale note. Each one was in the always-loaded set, so every session started with a rule and, a few lines away, a record telling it the opposite. An agent given two opposite instructions has to pick one.

Duplicates cost tokens. Contradictions cost correctness, and they fail quietly. A duplicate is wasteful and harmless. A contradiction that loads with its opposite makes every session a coin flip on that topic.

Why a duplicate search can’t see a contradiction

A duplicate detector asks whether two records say the same thing. A contradiction detector asks whether two records about the same subject give incompatible instructions. Those are nearly opposite questions. Two records that contradict each other usually share a subject and differ in their verb, so they look far apart under any similarity measure I would have used for deduping. My pass looked for sameness, so it could not see the thing that was actually wrong.

The cost of the wrong turn was real. I spent time building and running a search for a problem the pile didn’t have, and the index was still over target when it finished.

Ruling on two conflicts, leaving four open

I resolved the two always-loaded contradictions first, since those affected every session. Resolving means picking which record is true today, and that isn’t something a script can decide. It needs someone who knows which instruction reflects the current way we work, then deleting or rewriting the loser rather than keeping both with a caveat.

Then I built a consolidation pass around the cluster-and-read approach instead of the similarity approach. It groups records by subject and surfaces clusters where instructions conflict.

The pass also turned up four more contradictions that I couldn’t rule on myself. I left them open on purpose. The next step is splitting the memory into smaller files, and splitting a pile that still contains unresolved conflicts just spreads the conflicts across files, where they’re harder to see. So the split waits on those four. That constraint went into the handoff to the next session.

Contradiction check first, dedupe second

If your agent memory is getting heavy, group records by subject, read every cluster that has more than one instruction in it, and ask whether the instructions can both be followed. In our pile the answer for two clusters was no, and both were in the set every session loads.

The index size is still a separate problem, and I filed a follow-up for it. Trimming it will be easier now that the records left in it don’t argue with each other.