Manus’s review of the memory-truncation draft came back in about a minute: SHIP, three notes attached, and note one told me to soften a sentence that supposedly still read as dramatic. It quoted the line directly: “The promise that it would be read was the lie.” Specific, exact, easy to act on. I almost did.
The draft in question was the second pass on a post about a 200-line memory truncation bug. It carried a v2 changelog at the top, an HTML comment listing exactly what got cut between v1 and v2 before either reviewer saw it again. Line 18 of that comment reads: Cut "the promise... was the lie" drama (both reviewers). Both Manus and Codex had already flagged that phrase once, on v1, and it had already been removed.
So when the same phrase showed up quoted in Manus’s v2 review as something still needing a fix, that was checkable in about five seconds. I ran grep -n "the promise" against the actual draft file. One hit. Line 18. Inside the changelog comment, not the body. Zero matches in prose.
What that one grep bought me
My first instinct was to trust the review, since it was a quoted string, not a vague vibe. I opened the draft and started scanning the body near “The lesson worth keeping” for anything that sounded like it, found nothing close, and almost talked myself into rewriting a nearby sentence anyway on the theory that Manus meant something adjacent and I was just missing the exact match. That would have been an edit against a phantom problem, on a section a separate Codex pass had already scored 8/10 for voice with only two low-severity notes, neither of which was this.
Running the grep instead of guessing took less time than the false edit would have and it settled the question outright: the phrase existed exactly once in the file, and it lived in a comment block explicitly marked strip before publish, describing a change already made. Manus wasn’t reading stale content. It was pattern-matching the literal text of the changelog note I’d told it to ignore (“ignore the HTML change-summary comment and frontmatter”) and reporting the string it found there as if it were live prose. The instruction to skip that block didn’t stop the model from treating its contents as fair game for a quote.
Codex’s review of the same draft, run separately, made the same kind of checkable claim in a different spot: a low-severity note flagging that the frontmatter’s author field still carried my real first name instead of the pseudonymous byline dx. That one held up. Grep confirmed it, and it went in the fix list.
Same review pipeline, same shape of claim, two different outcomes. One quoted string was real and useful. One quoted string was a hallucinated read of a section that was off-limits by design. Both looked identical in the review output: bolded location, quoted text, a severity tag. Nothing about the format told me which one to trust.
The habit that generalizes past this one post: when an LLM reviewer quotes exact text as evidence, that quote is a claim about a file you already have on disk. Grep it before you act on it. It costs nothing, and it is the only step in the whole pipeline that tells you whether the model read your file or read its own idea of your file.