The daily alert said 17 tickets were inconsistent, and the repair for that exact condition was already written and ready to run. One command, and all 17 would be brought into line.
What “inconsistent” actually meant
The check compared each open ticket against a simple rule: certain fields should agree with certain other fields by the time work is considered done. Any ticket that failed the check got flagged, and the fix that already existed for a failed check was to update the ticket to match the rule, in bulk, across every flagged item.
Before running it, I went through the 17 by hand instead of trusting the flag. Sixteen of them were false alarms. The check inferred a ticket’s real state by reading a status log and guessing backward from it, instead of checking what had actually shipped, and it had no way to tell a ticket that had been formally retired years ago from one that had genuinely been delivered. Retired tickets and delivered tickets can both leave a status log that looks the same to a rule that isn’t checking the one thing that would actually distinguish them.
What the repair would have done
Ten of the sixteen false alarms were tickets from 2020 through 2022 that had already been retired, closed out and abandoned long before this check existed. Running the repair as designed would have stamped the current work cycle onto every one of them, rewriting a field on ten records nobody was working, based on a rule that never should have flagged them in the first place. That would not have been a cosmetic mistake. It would have been automated tooling confidently corrupting historical records because the detector feeding it couldn’t tell old and settled from new and broken.
What actually needed fixing
Only one of the seventeen was a real gap, one that genuinely disagreed with the rule and needed a hand correction, which I made directly rather than through the bulk repair. A second real gap turned up in the same pass, but it had already been flagged by the same watchdog weeks earlier and never acted on, so it wasn’t part of this batch of seventeen at all. The watchdog itself got the real fix: instead of inferring status from a log, it now checks actual delivery history, and it now distinguishes a retired ticket from a delivered one before flagging anything. One condition the original rule enforced is still open and unresolved on purpose, left for a deliberate decision rather than folded into this fix.
Why this generalizes past one ticket system
A detector and a bulk repair built for the same condition are usually built and reviewed together, as a matched pair, on the assumption that if the detector’s logic looks correct, the repair is safe to run on whatever it flags. That assumption only holds if the detector’s actual hit rate has been checked against ground truth, not just its logic reviewed on paper. A check that infers state indirectly, rather than reading the real source of truth, will eventually disagree with reality in a way that looks perfectly consistent from inside its own rule. The sixteen false alarms here didn’t look like bugs in the check. They looked like sixteen tickets that needed fixing, right up until someone checked what had actually happened to each one.
AI Skills
Use this lesson with the AI assistant you already use
A watchdog flagged 17 tickets as inconsistent and had a one-command fix ready to run. Sixteen of the seventeen were not actually broken.
Paste the prompt, share only the context needed to answer it, and treat the result as a draft for your review. Do not include confidential information or let an AI assistant make changes without your approval.
Optional: for a visual report and saved memory, run /dxdev first.
Don’t have it? Get it at dxdev.com/skills/dxdev. The prompt works without it.
dxdev LESSON · paste into your AI coding agent
LESSON: Before Auto-Repairing a Flagged Record, Prove the Flag Is Real
SOURCE: dxdev.com/blog/2026-08-18_the-repair-command-almost-ran
WHAT HAPPENED: A daily automated check compared open tickets against a rule about what "done" should look like, and flagged 17 as inconsistent. A one-command repair existed to bring flagged tickets into line automatically. Before running it, a manual pass through the 17 found that 16 were false alarms: the check inferred a ticket's real state from a status log instead of checking actual delivery history, and it could not tell a ticket that had been formally retired years earlier from one that had genuinely shipped. Running the repair as designed would have stamped the current sprint onto ten retired tickets from 2020 through 2022, rewriting history on records nobody was actively working. Only one of the seventeen was a real gap, one that genuinely disagreed with the rule and needed a hand correction, which I made directly rather than through the bulk repair. A second real gap turned up in the same pass, but it had already been flagged by the same watchdog weeks earlier and never acted on, so it wasn't part of this batch of seventeen at all.
THE RULE: An automated detector and an automated repair for the same condition should never be trusted as a matched pair without independently verifying the detector's hit rate first. A check that infers state from an indirect signal (a status field, a log line, a timestamp) rather than checking the actual source of truth (real commit or delivery history) will drift out of sync with reality, and a bulk repair built on top of that check will confidently apply the wrong fix at scale. Before wiring a detector to an automatic corrective action, run the detector alone, check its output against ground truth by hand, and only automate the fix once the false-positive rate is actually known, not assumed to be zero because the logic looked right.
CHECK MY CODE, then report PASS or FAIL with file:line for each:
1. Any automated detector wired directly to an automated repair/mutation action with no human review step in between, especially where the detector infers state indirectly rather than reading a primary source of truth.
2. Any check whose logic can't distinguish a formally closed/retired/archived record from an actively delivered one, if both can share the same superficial signal (e.g. a status field, a missing field, an old timestamp).
3. Any bulk-repair script whose blast radius (how many records it would touch, and how far back in time) was not measured and reviewed before its first real run.
THEN PRINT: a table (check, PASS/FAIL, evidence, fix) + a verdict (applies / partially / OUT_OF_SCOPE / no) + the single most important next action.