The release was one “safe to ship” vote away from landing with a save path that did not save what the UI claimed.

That was the uncomfortable result of a review cascade we have been using on release work. The cascade included our internal review, Manus, and Codex. On paper, it looked redundant. In practice, it produced two very different kinds of catches.

In the first review, the cascade found a missing server side permission check and an edit that was not transactional. In the second, most of the reviewers agreed that the change was safe to ship. The one reviewer that could read the actual files disagreed. It found a real save path bug that the others missed because they trusted an overconfident summary.

The difference was not model personality. It was not the number of opinions in the room. It was whether a reviewer could trace the behavior from the request through the code that persisted it.

A review can be unanimous and still be shallow

We do not run release review as a single approval gate. A change moves through an internal review and then through Manus and Codex. The point is not to ask several systems to restate the same description. It is to force competing reads of the change before we ship it.

That distinction matters because the input to a reviewer shapes the review. A summary can say that an edit saves correctly. A diff summary can say that permissions are handled. Both statements can be honest descriptions of intent. Neither proves the path is correct.

The release with the save path bug demonstrated the failure mode cleanly. The reviewers without file read access had a coherent account of the change. The account said that the save flow was implemented and safe. It was concise, confident, and wrong in the only way that mattered. The source reading reviewer followed the actual save path and found the bug.

Nothing about consensus repaired that missing evidence. Several reviewers can independently agree with the same persuasive summary. That is still one unverified premise copied several times.

This is easy to miss when a review system reports scores, approval language, or a pile of comments. A stack of “safe to ship” verdicts feels like strong evidence. It is not strong evidence if the people or agents issuing it are all looking at the same abstraction layer.

The first release had two different failure modes

The other release review made the same point from the opposite direction. This time the cascade caught two concrete problems: a missing server side permission check and a non transactional edit.

The permission problem was not a visual defect. The interface could present the right controls and the happy path could work. But a permission decision that exists only in the client is not a permission boundary. The server has to enforce the rule at the point where it accepts the change. Otherwise the application is relying on the screen rather than the system of record.

The non transactional edit was a different class of risk. An edit that touches related state needs an all or nothing outcome. Without a transaction, part of the work can persist while a later step fails. The user sees an edit attempt. The database can end up with an incomplete version of that edit.

Those two findings were useful precisely because they were not variations of the same complaint. One asked whether the server could reject an unauthorized request. The other asked whether the edit could leave related state half changed. The cascade produced value because it exposed both questions before release.

We considered a simpler workflow: let the internal reviewer handle the code and use Manus and Codex only for a final sanity check. We rejected that because it turns the outside reviews into a confidence ritual. They need a real chance to find something the first pass missed.

We also considered treating the consensus review as the final authority. The save path bug ruled that out. Consensus is useful evidence, but it is not a substitute for a reviewer who can inspect the source that creates the behavior.

The diagnostic path matters more than the verdict

A useful review result has to explain how it got there. “Safe to ship” does not tell us what was checked. “Bug found” is not much better if the reviewer cannot connect the bug to the code path that causes it.

For the missing permission check, the diagnostic question was direct: where does the server decide whether this request is allowed? If the answer is only a client control, the review is not finished.

For the non transactional edit, the question was equally concrete: which persistence steps belong to one operation, and what happens if a later step fails? If there is no transaction around related writes, then the edit has an incomplete state waiting to happen.

For the save path bug, the question was simpler still: what file receives the saved value? The summary supplied an answer. The source supplied a different answer. Only one of those answers was running in the release.

That sequence is the part I want to preserve. We observed a confident approval. We compared it with a review that had direct file read access. We traced the claimed behavior through the actual save path. The disagreement was not subjective. The code resolved it.

This is why I am wary of review systems that optimize for polished conclusions. An overconfident summary is dangerous because it removes the reviewer’s urge to verify. The cleaner the summary sounds, the easier it is to mistake it for an implementation trace.

File read access changes the job

Giving a reviewer source access does not mean the reviewer will always be right. It means the reviewer has the ability to test a claim against the implementation rather than against a retelling of the implementation.

That is a different job. A summary reviewer asks whether the described approach sounds reasonable. A source reviewer asks whether the route, permission boundary, edit sequence, and save path actually match the description. Both can be valuable. They should not receive the same weight.

The practical change for us is simple. We still run the internal plus Manus plus Codex cascade. We still want independent review because the permission and transaction findings show what multiple passes can uncover. But every release review now needs at least one reviewer with real file read access, and that reviewer’s diagnostic path has to be visible.

A release is not safer because four reviewers agree. It is safer when one of them has read the code that will run, traced the behavior that matters, and can show the rest of us where the summary stopped matching the source.