---
title: "Infrastructure You Thought Was Running Isn't"
canonical: https://dxdev.com/blog/2026-08-30_guard-dead-everywhere/
datePublished: 2026-08-30
---
A deleted line in a colleague's branch would have wiped a team's Stripe keys the next time they saved their payment options. Not flagged by CI. Not flagged by a linter. Caught because I read the diff by hand.

A colleague had rebuilt the fundraiser admin panel and pushed it up for review. I went through it line by line and found three separate defects sitting on the branch, already pushed, already past whatever automated checks were supposed to run: the deleted line that would silently wipe a team's stored Stripe keys on their next save, an Active toggle that never reverted in the UI when its save call failed (so the screen would keep showing "on" after the backend had already said no), and group names flowing straight into link-dialog HTML unescaped, which is XSS-in-waiting the moment someone names a group `<script>`.

Every one of those is exactly the category of thing a pre-push guard exists to catch. We have one. It's a hook that runs before code leaves a developer's machine, and it's supposed to run on every clone of a large legacy codebase. So the question wasn't just "how did three bugs get through," it was "why didn't the thing whose entire job is stopping this ever fire."

## Checking whether the guard actually ran

I checked that colleague's clone first, since that was the branch in front of me. The guard wasn't running there. My first assumption was a local misconfiguration on their machine, maybe a missing hook install step from when they'd cloned the repo. So I checked mine. Also dead. I checked a third clone. Same result. That ruled out "one developer's setup is stale" and turned it into "the guard has not been protecting anyone, on any machine, for however long this has been true."

That's a different kind of bug than the one I walked in expecting to file. I filed a ticket for it: the pre-push guard is dead in every clone of the codebase. Not flaky, not intermittent. Dead everywhere it was supposed to be running.

While I was in there I also found the review skill I'd been using to triage tickets was assuming the wrong developer on attribution, a separate bug but discovered in the same pass, so I fixed that too before moving on.

## What "dead" actually meant

The failure mode is the part that matters more than the discovery. A pre-push hook that blocks a push loudly is annoying but safe: you know immediately that something stopped you, and you go find out why. A pre-push hook that silently declines to run and lets the push through anyway is worse than no hook at all, because it gives you the confidence of a safety net without the net. Everyone who touched that codebase, myself included, had been pushing under the assumption that a set of checks was standing between "code I wrote" and "code in the shared branch." None of it was.

We didn't get a full trace of why the guard had gone quiet on every clone at once, and I'm not going to overstate the root cause here. What we do know, from what shipped after: the fix that landed wasn't "make the guard more robust so it never crashes." It was "when the guard itself crashes, make that fact loud instead of silent." The follow-up commit title says it plainly: the pre-push guard degrades open when the guard itself crashes. Before that, a guard that threw an exception on startup could fail in a way indistinguishable from success, no error, no warning, the push just goes through. After that fix, a guard crash still lets the push through, deliberately, because blocking every developer's work over a bug in the safety check itself is its own outage. But now it says so. You get a visible warning that the check didn't run, instead of silence that reads as a pass.

That's the actual tradeoff we made: fail open, but never fail quiet. Failing closed on guard crashes was the other option on the table and we rejected it, because a bug in the guard shouldn't get to hold every developer's pushes hostage until someone notices and fixes the guard itself. But failing open silently is what actually happened here, and it's what let three real defects, one of them a data-loss bug against live Stripe keys, sit on a pushed branch with nobody, human or automated, having actually looked.

## The part that isn't really about the guard

The bugs got caught. That's the good half of this story: manual review is still the backstop, and it worked. But "manual review caught it" is not a plan, it's what you fall back to when the automation you were counting on turns out to have quietly stopped counting. The guard existing in the repo, referenced in onboarding docs, assumed by every developer who'd been told it was there, did none of the work its presence implied. The gap between "the file is in the hooks folder" and "the check actually runs on every push" was invisible until someone happened to read a diff closely enough to need to ask why it hadn't been.

If you're relying on a pre-push or pre-commit guard anywhere in your own stack, the useful test isn't "does it exist." It's "when did I last watch it actually fire and block something." If you can't answer that, you don't know whether you have a guard or a folder with a script in it that nothing calls.
