---
title: "Fixing Code You Can't Reproduce"
canonical: https://dxdev.com/blog/2026-08-22_targeted-hardening-without-repro/
datePublished: 2026-06-25
---
At 3:36 PM, we shipped a release because one organization-account tournament event was stuck and the exact trigger would not come back on demand.

That last part matters. Reproduction is the gold standard for a bug report. It tells you what happened, lets you prove the fix, and gives you a test that guards against the next regression. I want it whenever I can get it.

But “cannot reproduce” is not the same as “nothing can be fixed.” It can mean the failure depended on state we no longer had, an intermittent sequence, or a route through code that should not be available in the first place. In this case, the useful work was not manufacturing confidence about the original trigger. It was removing the reachable condition that could produce the kind of state we had just seen.

## A stuck event is evidence, not a complete explanation

The reported failure was concrete. A tournament event on an organization account was stuck. We could not recreate the exact action sequence that put it there.

That left two separate questions, and treating them as one would have slowed us down.

The first question was forensic: what precise sequence caused this individual event to get stuck? We did not have a reliable answer. The report did not turn into a repeatable local case, and I was not going to invent one in the ticket just to make the story tidier.

The second question was architectural: what state or mode is still reachable in the organization-account event path that should no longer be reachable? That question did have an answer. The deprecated preview mode could still be forced onto an organization-account event.

I do not know whether that was the exact historical trigger for the reported event. The source record does not establish that, and the distinction is important. What it establishes is that the path was still open, it was deprecated, and closing it protected the affected account class either way.

That is enough to act when the change is narrow and the failure mode is bad. It is also a real bet: I signed off on a hotfix without the thing I usually want most, and I have to live with the possibility that the actual cause is still out there, untouched, waiting for the next stuck event on an organization account to look identical to this one and prove me wrong. I would rather ship a defensible constraint I can explain than sit on a comfortable silence while the deprecated path stays open.

## The fix was to make an invalid route unavailable

The production change hardened the organization-account event path so that events can no longer be forced into deprecated preview mode.

This was not a broad rewrite of tournament setup. It was a constraint at the boundary where the old mode could still be selected or applied. The objective was simple: an organization-account event must not enter that mode, regardless of what input or stale state tries to push it there.

That is a materially better outcome than a patch that only repairs the reported event. A one-record repair might relieve the immediate incident, but it leaves the system capable of making the same bad choice again. Blocking the bad state at the code path turns a conditional warning into an invariant.

> An invariant is not “we expect this not to happen.” It is “this class of object cannot enter this state through this path.”

The hardening was also intentionally targeted. We had one stuck event, an unreproducible trigger, and a deprecated mode that should not have been reachable for this account class. That did not justify rewriting every tournament-event transition or making a speculative change across unrelated account paths. It justified closing the door that should already have been shut.

## The discarded path was an endless reproduction hunt

I could have kept the incident open until we had the exact clicks, timing, and data state. That is the cleanest debugging narrative, but it was the wrong release criterion here.

A perfect reproduction would have told us more. It might have identified the original sequence. It might have given us a highly specific regression test. But waiting for one would have kept a known-deprecated mode reachable while we chased an intermittent condition that we could not observe again.

The other bad option was to treat the event as an isolated data problem. That would make the current incident disappear without changing the condition that made it possible. The fact that the report involved one event did not mean the defect belonged to one event.

The decision was not “fix without evidence.” The evidence was the stuck event plus an accessible deprecated mode in the affected code path. The decision was to fix the evidence we could prove, and be precise about the part we could not.

## What we documented separately

The production hotfix did not become a bucket for every tournament problem observed that day. A separate intermittent blank single-competition setup page went into its own item.

That separation is part of the engineering work. A vague report can tempt you to join failures because they happen near the same feature. Doing that turns diagnosis into narrative. The stuck event was addressed by preventing deprecated preview mode on organization accounts. The blank setup page remained its own intermittent problem until it has its own mechanism and fix.

The hotfix took about 2 hours and 30 minutes from investigation through shipment. At the end of that time, we still did not have a faithful reproduction of the original report. We did have a narrower system: an organization-account event could no longer be forced into a mode we had already deprecated.

I would rather ship that honest constraint than wait for a perfect story. Reproduction proves a cause. Hardening removes an opportunity for failure. When the old path should not exist, removing it is not a compromise. It is the fix.
