---
title: "The Alarm That Trained Everyone to Stop Reading It"
canonical: https://dxdev.com/blog/2026-09-10_the-alarm-that-trained-everyone-to-stop-reading-it/
datePublished: 2026-09-10
---
There was a bot whose whole job was noticing and quietly fixing small problems in a piece of engineering tooling, and it kept crashing. Every time, it sent the same kind of alert: a crash message and a timestamp, several times over, the moment it happened. After enough of those, you stop reading them the way you'd read the first one. That's the real cost of an alert that fires without being useful. It doesn't just interrupt you. It trains you to ignore the next one, including the one that actually mattered.

I set aside an afternoon and pulled up the actual causes instead of reacting to the symptom each time.

## Three ordinary things, wearing different disguises

The first cause was a leftover marker from an earlier step that had died partway through and never cleaned up after itself. Every attempt after that hit the same leftover marker and refused to proceed, over and over, until someone went in and cleared it by hand.

The second was a set of routine internal checks that were failing, not because anything was actually broken, but because the machine was busy with something unrelated at the same moment the check ran. A check that fails once under normal load and gets treated as a hard failure produces an alert that says "something is broken" when what actually happened is "the system was busy." So those checks got a retry before they were allowed to escalate into anything.

The third was the simplest and least interesting: the process itself sometimes died and just stayed dead until a person noticed and restarted it. That one didn't need a diagnosis. It needed to restart itself.

## The fix that looked done and wasn't

The first fix, for the leftover marker, was built the right way on paper: before clearing anything, check whether something is genuinely still using it right now, and only clear it if that check comes back clearly clean. Refuse to guess.

It kept not firing anyway, on more than one separate occasion, on exactly the cases it was built to protect. The reason took a while to find: the check that confirmed nothing was still using the marker relied on a single outside query, and that query itself was timing out under the exact kind of load that causes the leftover marker to appear in the first place. When the query timed out, the check correctly refused to guess and left everything alone, which sounds safe in isolation. In practice it meant the fix's own safety net went blind at precisely the moment the underlying problem was most likely to be happening. A cautious check that can't get an answer under real load isn't cautious. It's quietly useless right when it's needed most.

The actual fix for the fix added a second, cheaper way to check when the first one couldn't answer, plus one narrower rule for when both stayed unclear: if the leftover marker was empty and old enough that nothing legitimate could still be using it, that counted as a second, different kind of evidence, not a guess.

## Where the routine stuff goes now

Alongside fixing the mechanism, the bigger change was deciding where routine, self-resolved anomalies go. If the process restarts itself and comes back clean, if a check fails once and passes on retry, if a leftover marker gets cleared automatically, none of that needs to interrupt anyone. It gets logged and rolled into a single daily summary instead. What still pages right away is only the thing the automatic layer actually tried and failed to resolve, which turns out to be a much smaller and more honest list than "anything that looked wrong at any point."

## Where AI fits

An assistant is genuinely useful for the grouping work: looking at a pile of repeating alerts and sorting them by actual cause instead of message text, then proposing a safety check that only acts on a confirmed answer. It shouldn't be the one deciding, on its own, that a stalled or inconclusive check is close enough to a clean one, and it shouldn't quietly stop paging someone without that change being visible somewhere a person can see it.

## The human decision

A person decides which failures are genuinely safe to auto-resolve and log for later, and which ones must always interrupt someone immediately, no matter how routine they start to look. That's a judgment about real risk, and it doesn't belong to the system fixing itself.

## The lesson

A repeating alert almost never means a new problem. It usually means the same small handful of ordinary causes, wearing different disguises, and a fix that only gets tested against a quiet, idle system will look done and then quietly fail on exactly the busy day it was built for. Fix the actual mechanism, not the symptom, and then go check that the fix still answers when things are genuinely loaded, not just when they're calm.

The paired Build Log walks through all three causes and the exact reason the first fix stayed silent under load.
