The morning the alarm cried wolf fifty-five times
We built a watchdog to catch a real, known problem before it became a crisis again. On its first real morning of actual use, it paged 55 times, all at once, all saying the same thing: stuck, broken, needs attention.
Forty-four of the 55 weren’t broken at all. Each one was a site still waiting on a routine migration step, sitting in a state that looked identical to a real failure if you only checked the one thing the watchdog was checking.
Why one signal wasn’t enough
The watchdog looked at a single piece of recorded state to decide whether something was healthy. That single signal is genuinely useful, most of the time. It just cannot tell the difference between “this hasn’t happened yet because it’s not supposed to yet” and “this is supposed to have happened and didn’t.” Both look exactly the same from that one angle.
A pile of 44 false alarms would have been an annoying but survivable morning on its own. What made it a real problem was the other eleven, sitting inside that same batch of 55, wearing the identical “stuck” label: sites whose secure address returned a connection failure while their plain address still answered, meaning they were genuinely down for anyone trying to reach them securely. One signal couldn’t tell those eleven real outages apart from the 44 non-events, because from where it was looking, all 55 looked exactly the same.
What changed
The fix added a second, independent signal: instead of trusting the recorded state alone, the watchdog checked where the thing actually resolves to right now. That second check split the batch correctly. The 44 that were genuinely fine got logged, counted, and muted, with that muted count printed plainly so a quiet morning could never be confused with a clean one. The eleven that were actually broken got paged.
One more piece mattered just as much: before paging anyone, the watchdog now re-checks its finding against fresh, current data one more time. A finding from even a few minutes ago can already be stale. Paging on a stale finding wastes a person’s attention on something that may have already resolved itself.
What a person still has to decide
None of this makes the watchdog the one deciding what’s an emergency. It decides what’s worth putting in front of a person. A person still reads the page, decides how urgent it really is, and decides what to do about it. What the second signal buys is confidence that when the page arrives, it’s pointing at something real, not diluted inside forty-four false ones that would have trained anyone to stop reading it closely.
The rule worth keeping
If your monitoring only checks one thing, it will eventually page you for something that’s fine and stay quiet, hidden in plain sight, about something that isn’t. Before you trust any alert to page a person, ask what it would take to fool that alert with something completely healthy. If the answer is “nothing, one field is enough,” add a second, independent check before the page goes out.