---
title: "34/34 Green Tests, Then I Made a Second AI Argue Against My Own Suite"
canonical: https://dxdev.com/blog/2026-08-22_adversarial-verifier-before-touching-prod/
datePublished: 2026-06-17
---
Thirty-four tests were green when I stopped trusting them.

We were preparing new safety-net code for a live SMS integration. The queue already had idempotency, a kill switch, retry behavior, and stalled-job recovery. The change closed the window where a recipient is accepted for sending but the job is not yet marked sent. That is a small window until a process crashes, a worker stalls, or a provider rejects one recipient while others are in flight. In a live queue, the edges are the system.

I wrote the tests first. Thirty-four passed. I had a passing suite, a change set that looked disciplined, and all the usual reasons to call the work finished.

I did not touch production. I gave the passing suite to a second model and asked it to attack the tests rather than admire them.

## The green suite was only the first gate

The original implementation had three requirements. It had to record an accepted message before a crash could create an avoidable duplicate send, stop if the kill switch engaged mid-job, and distinguish a retryable provider failure from a hard error such as `21211`.

The first suite covered the happy sequence and obvious fault paths. It exercised retries, kill-switch behavior, and a simulated crash. It went green at 34 of 34.

The count showed the code met cases I had anticipated. It did not show that I had anticipated the right cases. A green suite can hide a wrong model of time, an incomplete classification, or a safety check that never runs in the worker that matters.

So I made the verification adversarial. The second model got the diff and the tests, with one responsibility: find the path my passing suite had not made impossible. It checked whether the tested behavior defended the queue at its failure boundaries.

## The gap was in the clock, not the branch

The verifier found a problem in the test double for Redis. Our fake implementation did not model time-based TTL expiry like the real store. That mattered because the safety net uses a short-lived key to control recovery behavior after an interrupted send. A crash test that only manipulates values, without advancing the clock, can look convincing while never proving that the key expires under the condition the production queue will see.

The diagnostic path was straightforward once the verifier pointed at it. The crash test appeared to test expiration because it created the short-TTL state and then ran recovery. But the fake store had not moved forward in time. The test was asserting a sequence against a static key-value model, not a sequence against expiry.

I changed `fakeRedis` to model real time-based TTL expiration and added `_advance` so the crash test could drive the clock deliberately. That moved the test from "a key was present when recovery ran" to "the key expired after the configured interval, then recovery took the expected loss-safe path." It also forced a configuration rule into the code: `SHORT_TTL_MS < stallInterval`.

That invariant is not decoration. If the short TTL is allowed to outlive the stalled-job interval, the recovery logic can make a loss-safety promise it cannot keep. I added `assertTtlInvariant()` to enforce the relationship instead of relying on a comment or a remembered deployment setting.

## The retry was necessary, not sufficient

The other hard edge was the gap between accepting a recipient and marking it sent. I narrowed it by retrying `markSent` three times. If all three attempts fail, the worker emits a `mark_sent_failed` alert. That is not an attempt to pretend the duplicate window disappears. It is a decision to accept a narrow residual duplicate window, monitor it, and make the unrecorded state visible.

We considered treating every failure around `markSent` as retryable. That lost because a provider hard error needs a different outcome. Error `21211`, for example, is classified as non-retryable when it is thrown in the catch path. Retrying it only increases noise and keeps a bad recipient in a cycle that will not resolve.

We also considered checking the kill switch only once, before the job begins. That lost because an operator can engage it after the first recipient. The final path checks it per recipient, so a mid-job switch stops the rest of the batch rather than waiting for the next job.

Those choices made the code more explicit about its limits. The `markSent` retries reduce one dangerous window. The alert exposes the residual one. The per-recipient kill switch makes an operator action take effect in the middle of execution. The non-retryable classification keeps a known bad destination from looking like transient infrastructure trouble.

The verifier also found two comments that had become stale after the changes. That is not the headline failure, but I fixed them. In queue code, stale comments are dangerous because they tell the next person a safety boundary is somewhere it no longer is.

## Forty cycles before the handoff

After the fixes, the suite reached 39 of 39 and ran cleanly 20 times. I sent the revised result back through verification. The second pass confirmed the chosen design, the narrow duplicate window remains accepted and monitored, rather than falsely eliminated.

There is still one production-protection pass left. `worker.js` is not yet wired to use these modules, and startup must call `assertTtlInvariant()`. Until both are in place, the hardening lives in tested code but not in the live worker path. That distinction is why I stopped at the boundary instead of treating 39 green tests as permission to cut over.

The useful move was not asking a model to write more code. It was making a second model argue with the suite I was already proud of. The first 34 tests told me the implementation matched my plan. The verifier found the place where my plan had failed to model production time. The final 39 tests are better because one of them is a test I would have shipped without.
