---
title: "The Safety Check I Trusted Wasn't One"
canonical: https://dxdev.com/blog/2026-07-24_the-safety-check-i-trusted-wasnt-one/
datePublished: 2026-07-24
---
I had just finished a cleanup on a live database: a gate to stop a background scanner from creating ownerless records, plus a one-time pass to clean up the ones that had already piled up. Standard discipline says a change like that gets a paired rollback script, written and staged before the real thing runs, in case any part of it needs undoing later. I wrote one. Then I made a mistake trying to prove the rollback script itself was safe.

## A safety check that wasn't checking anything

Before running the rollback script for real, I wanted to confirm its syntax was correct without actually executing it, so I switched on a mode built for exactly that: parse the statements, don't run them. I ran it against the live database, treating that as a harmless no-op. It is a no-op, for statements inside a stored procedure body. It is not a no-op for plain schema and data statements sitting outside one, and most of a rollback script is exactly that: drop this constraint, drop this table, restore these rows.

So the "safety check" ran the script. For real. A constraint I'd just added got dropped. A small reference table the cleanup depended on got dropped with it. The rows I'd just finished cleaning up reverted to their messy prior state, as if the cleanup pass had never run.

## Catching it in the same breath

The only reason this stayed a non-event is that I was already running verification queries as part of the same validation pass, and they came back wrong immediately: a constraint that should exist, didn't; a table that should have rows in it, didn't exist at all. There was no gap between causing it and noticing it. I restored the constraint with its correct definition, recreated and reseeded the reference table, ran the cleanup step again, since it was written to be safe to repeat, and checked the end state against what the cleanup was supposed to produce. It matched, with one exception: the audit log now carries a duplicate pair of removal entries for those rows, left in place because that log is never edited after the fact. Nothing about this ever reached anyone using the actual product; the whole arc happened inside a validation step meant to catch problems before they could.

## The bug this exposed

While putting the pieces back, I found a second problem sitting in the same rollback script, one that hadn't fired yet because nothing had needed the rollback for real. The step that re-narrows a constraint after undoing a change assumed the underlying log table would only ever contain the values it had at the start. But that log is append-only, records get added, never removed, and the gate I'd shipped that same day had already logged a new kind of entry. Run the rollback as written, and that step would fail outright the moment it hit a single one of those new rows. I fixed it to allow existing rows through without re-validating them, which is the correct behavior for an append-only log, and confirmed the corrected version handles both the untouched case and the case where new rows already exist.

## What the mistake actually was

I didn't skip a safety step. I added one, on purpose, specifically to avoid running something risky against a live system before I was sure it was correct. The mistake was trusting what the feature's name implied instead of checking what it actually covers. "Parse without executing" sounds like a blanket guarantee. It is a guarantee about one specific case, stored procedure bodies, and silent about everything else, including the exact kind of statement a rollback script is mostly made of.

The fix isn't "don't use dry-run modes." It's that a mode's name is a claim, and a claim about safety is exactly the kind of thing worth testing on something that isn't live before trusting it on something that is. I had the discipline to write a rollback script and to try to validate it before use. I didn't have the discipline to check that my validation method actually validated anything, and against a shared production system, that gap is where the real risk always hides, not in skipping the safety step, but in trusting the wrong one.
