---
title: "34 of 34 Tests Passed, So I Stopped Trusting Them"
canonical: https://dxdev.com/blog/2026-06-17_all-green-doesnt-mean-all-tested/
datePublished: 2026-06-17
---
# 34 of 34 Tests Passed, So I Stopped Trusting Them

We were changing safety-net code sitting under a live messaging queue, closing a narrow window where a message could get accepted for sending but not yet marked as sent. Miss that window and a crash at the wrong moment could mean the same message going out twice. Get it right and the fix is invisible, which is exactly the kind of change where "it looks fine" is the least reassuring thing you can say about it.

I wrote the tests before writing the fix, the way you're supposed to. Thirty-four of them, covering retries, a manual kill switch, and a simulated crash partway through. Every one passed. That's usually the point where a piece of work gets called finished, because a full green suite feels like proof.

Instead of shipping on that feeling, the passing suite and the change itself went to a second, independent reviewer with one specific instruction: don't check whether this looks reasonable, try to find whatever the passing tests failed to rule out. Not a second opinion looking for reassurance. An assignment to actually attack the work.

That reviewer found something the original thirty-four tests couldn't have caught, because the gap wasn't in the code being tested. It was in the tests themselves. The fake version of the underlying storage system used to simulate a crash didn't actually model how time-based expiration really behaves. A test could look completely convincing, crash simulated, recovery triggered, verdict green, while never once proving that a value would genuinely expire under real timing conditions the live system would actually face. The suite wasn't lying. It was answering a slightly different question than the one that mattered, and nothing about reading it more carefully would have revealed that, because the blind spot lived one layer beneath the tests, not inside them.

Finding that also surfaced a second thing worth fixing properly instead of just noting: a specific timing relationship, one safety window needing to stay shorter than another recovery window, had been true by luck rather than by design. That got turned into an active check that would fail loudly if it ever stopped being true, instead of staying something only a comment reminded you about.

## If your own tests just went fully green

That result tells you your code does what you designed it to do, against the situations you thought to check. It says nothing about the situations you didn't think to check, and it says nothing at all about whether the tools simulating a real failure actually behave like the real thing. A second reviewer whose whole job is finding what the passing suite didn't cover, not agreeing that it looks solid, is often the only way to find that kind of gap before a customer does.
