---
title: "The Report Said We Fixed 13 Problems. We Hadn't."
canonical: https://dxdev.com/blog/2026-07-11_the-report-said-fixed-it-hadnt/
datePublished: 2026-07-11
---
A recurring-failure report read 46. The very next run of the same report read 33. Nothing in the actual work had improved in between. The counter itself had been wrong.

These reports exist to answer one narrow, useful question: what did we actually have to fix, because the same problem kept coming back? This time, the tool answered a different question without saying so. It was counting routine maintenance edits as if they were fixes for recurring problems.

That made the report worse than useless. It handed over a precise, confident-looking number, while quietly mixing ordinary upkeep in with the real failures it was supposed to be tracking.

## A plausible number is not the same as a correct one

The mistake did not show up as an error message or a blank result. The tool ran fine. It found changes, it added them up, and it produced a total that even looked like a real result, large enough to suggest something was going wrong repeatedly.

But a routine maintenance edit is not automatically evidence of a recurring problem being fixed. It can just be ordinary upkeep of supporting material that happens to sit near the real work. Treating it the same way turns the metric into a count of activity, not a count of learning. A count of 46 was really a count of 46 things that happened to satisfy a bad definition.

## Checking the tool before trusting the result

The actual fixing pass had still found real recurring problems worth correcting. That part of the work was valid. But the reported total did not line up cleanly with what could actually be accounted for, which is what raised the question in the first place.

The easy option was to leave the number alone and treat it as a rough, directional signal. That would have quietly given up the entire point of tracking it. A noisy total can't answer whether a problem is actually returning less often, or whether a batch of unrelated maintenance just happened to move the number.

Adjusting the count after the fact, or lowering the bar so the number looked more comfortable, wouldn't have fixed anything either. It would have just made a wrong number feel better. The actual fix had to happen at the point where the tool decided what counted in the first place, separating a genuine correction from maintenance that only looked similar on the surface.

## A less flattering number that was actually useful

Once that fix landed, the same report read 33. A drop from 46 to 33 sounds, on its face, like thirteen wins. Reading it that way here would have repeated the exact mistake in a different form. Nothing was newly fixed between those two numbers. Thirteen items that were never real problems stopped being counted as if they were.

That makes the remaining 33 worth more, not less. They're closer to the actual set of things worth acting on. A measurement bug inside a tracking system is real work in its own right, because it changes what you believe is true about everything else the system reports.

## The rule

A number that looks precise is not the same thing as evidence. Before trusting a metric to tell you whether something is improving, check what each item inside that count actually is, because a tool that quietly redefines what it's counting can hand you a confident, wrong story about your own progress.
