At 11:06 PM on June 29, I had a post whose core failure was a 200-line or 25KB context cap, and I still did not let it ship after one AI said SHIP.
The failure itself was ordinary enough. I had a long-lived agent whose memory file was intact on disk, but its startup harness loaded only the first 200 lines or 25KB, whichever boundary it hit first. New memories were appended at the end, which meant the newest entries were precisely the ones the model never saw. The fix was a small always-loaded policy file, a separate lookup catalog, and a guard that exits non-zero on overflow or an orphaned entry.
The writing risk was factual softening. A draft can turn “the harness loads the first 200 lines or 25KB” into “the agent kept forgetting,” then leave the reader with a model-behavior explanation for an architecture failure. That was close enough for me to change how a post reaches publish.
The sentence I would not let us blur
An early draft had a strong but loose line about the memory file “lying.” It also described the agent as forgetful. Both phrases carried the feeling of the bug, but neither named the mechanism. The important fact was simpler: the file was correct, the startup context was incomplete, and the append order put the newest entries beyond a silent boundary.
That distinction became the test case for the review pipeline. I do not send a post to several agents and count favorable opinions. Each reviewer gets a different job, a fixed verdict, and the artifact it needs to inspect. A ship verdict does not cancel a later revision request.
The first pass went to Manus. It returned SHIP, with three concrete edits: remove the theatrical “worst class of bug” paragraph, replace the rhetoric about a broken promise, and state the portable idea as “separate state from context.”
Codex then returned REVISE. Its strongest objection was technical. “Everything past that line never reaches the model” was imprecise because the cap had two dimensions, lines and bytes. It also pushed the post to distinguish an always-loaded policy from a lookup store. That is the architecture. “The model was fine” was too absolute, because the failure was architectural without proving every possible model behavior was fine.
The reviews disagreed on the verdict and agreed on several changes. That is why I want both. A single reviewer can polish an inaccurate framing while making it more persuasive.
Seven lenses, not one blended score
After the first revision, I run the body through a seven-dimension sweep. Fact and security or privacy both passed at 10. Humanization, voice, publishability, and subject passed at 8, 7, 8, and 9. Craft returned REVISE at 6.
The point is not to pretend a score makes prose true. It is to keep fact checking, identity risk, readability, and craft from collapsing into one vague “looks good” decision. In this case, craft found a paragraph that restated the fix after the fix had already done its work. It also found a three-sentence closing that was trying to explain the idea a second time. I cut both down.
The sweep created a real conflict. Manus and Codex wanted the generalization, “separate state from context,” while the craft reviewer wanted the whole lesson paragraph removed as redundant. I kept the generalization and removed the restating tail. The subject reviewer had scored portability at 9, while craft objected to the duplicated explanation, not the concept.
The final pass reads the artifact, not the mood
The adversarial verifier gets the revised post and the implementation artifact. Its job is to look for drift between them, not to praise the prose.
In this case, it found three legitimate changes after the earlier passes had finished. First, it removed a dead transition before the fix section. Second, it required the guard to be concrete: a short script checks line count and byte size against the cap. Third, it rewrote one self-check instruction so a manual diff could not be mistaken for something the guard performed automatically.
Then it checked the post’s structured claims against memory_check.py. The 200-line and 25KB boundary, append plus head truncation, the two-tier design, and the non-zero exit for overflow or orphans all matched the code. That is a different review task from “does this read well?” It is closer to a test attached to an essay.
The release path is intentionally inconvenient
draft → Manus: SHIP or REVISE → Codex: SHIP or REVISE → revision and seven-dimension sweep → targeted synthesis of conflicts → adversarial verifier against the source artifact → publishI considered three easier versions. One general review pass lost because one agent can polish an inaccurate framing. Voting lost because Manus shipped while Codex found precision problems, and the craft dimension still required changes. One giant rubric prompt lost because the verifier needs the real code or config beside the prose, not a longer opinion about the prose.
The pipeline has a cost. It adds revisions, keeps a good draft from going live on the first encouraging verdict, and makes me resolve conflicts instead of pretending the agents agree. It also keeps a memory-cap article from becoming a generic story about unreliable AI.
The 200-line and 25KB cap was never the story; the append order was. A memory file that only grows forward guarantees its most recent entry is also its most likely casualty, and no amount of favorable verdicts from Manus or Codex changes that arithmetic. What changed the post was the same thing that changed the harness: a second, always-loaded file that does not wait its turn at the end of a growing log.