---
title: "The Review Said Clean. A Second Reader Disagreed."
canonical: https://dxdev.com/blog/2026-06-25_the-review-that-said-clean/
datePublished: 2026-06-25
---
The review came back clean. Nothing was exploitable, every change was scoped to the right place, and the one real risk it did flag already had a fix proposed. I agreed with all of it, and sent the same work to a second, independent reviewer anyway, mostly out of habit.

## Two different questions

The first review had asked one question well: does this pass the checks we already know to run for. It did. What it hadn't asked was a second, harder question: what happens on every input a real person, using the tool the ordinary way, could actually produce. Those are not the same question, and the gap between them is exactly where the second reviewer spent its time.

## What a checklist is built to see

The one issue the first review caught was a real one, rated correctly, with a proposed fix already on the table. What it missed was whether that same risky pattern existed anywhere else doing conceptually the same job. It did, at a second location the proposed fix never touched, protected only by a comment that described behavior the code no longer matched.

The rest of what the second pass found didn't look like a bug at all on the surface. One operation, given an empty starting value it wasn't expecting, didn't throw an error. It silently produced an empty result and saved it, replacing something real with nothing. Another treated a blank filter field as if it meant "match nothing" instead of "match everything," quietly narrowing a bulk change to a tiny sliver of what was intended, while the success message still reported the number that would have applied to the whole thing.

None of that trips an error page. Every one of them looks, from the outside, like the operation worked.

## The pattern behind every one of them

A checklist built from known failure modes is good at catching loud, structural problems: the kind that throws an exception, the kind that lets someone see data they shouldn't. It is not naturally built to catch a write that succeeds, reports success, and is simply wrong. That is a different shape of failure, and it needs a different question, not a longer checklist.

## Why the second read still mattered

It would have been reasonable to treat one clean review as enough and move on. What made the second pass worth running was not doubt about the first reviewer's competence. It was recognizing that "does this pass the checks we built" and "what can a real user actually make this do" are structurally different questions, and no single pass answers both by accident.

## Where AI fit, and where it stopped

Reading the same code with the second question in mind, and working through what a blank field, an empty value, or a second call site would actually do, is work an AI reviewer can do quickly and show its reasoning for. It found the pattern, not because it was smarter than the first review, but because it was asked a different thing.

Deciding that this change was worth a second, independent look in the first place was a judgment call a person made. Deciding what to do about the rows the quiet mistakes had already touched, and whether anything needed to be corrected for anyone already using the feature, stayed a human decision too. An AI agent can tell you a write is silently wrong. It should not be the one deciding what happens next for the people it already touched.

## The lesson

The Build Log companion lists everything the second pass turned up. The rule that came out of it: a review that answers "does this pass what we already check for" is not the same as a review that answers "what happens on every input a real person could produce," and the second question is the one that catches a write that succeeds, reports success, and is simply wrong. Change the trigger for a second look from "does this touch something sensitive" to "does this write based on a selection a person built by hand," and you will catch that quiet kind of wrong before a customer does.
