A monitor we run watches every domain we serve for certificate and connection trouble, and pages the most urgent kind as a live outage. One customer’s domain had triggered that urgent page on six of the previous nine days. Someone finally asked why that particular domain kept getting the serious warning instead of the routine kind.

What I believed when I was asked

I had an answer ready, because the alert’s logic is simple: it fires when a hostname is not being served through our edge and the server it does point to cannot complete a basic secure connection. This domain met both conditions exactly, so I explained the classification as correct. It read like a good answer. It restated the rule and confirmed the rule had fired the way it was built to fire.

What was actually true

That answer never touched the actual question, which was not “did the alert’s condition match” but “is this outage ours to act on.” It was not. The domain in question pointed at a server the customer runs themselves, somewhere else entirely, and that server being down had nothing to do with anything we operate. The alert had been built for a narrower, real problem: a migration cleanup had retired some of our own server configuration while customer records still pointed at it, so a handful of sites lost secure access even though the older, less secure version kept working. That is genuinely urgent, because it is ours. A customer’s own third-party host being unreachable is not the same category of problem at all, and treating it as one meant six false urgent pages in nine days over something no fix on our end could ever touch.

Getting there took two passes, not one. The first time, I defended the classification because the trigger condition was, in the narrowest sense, true. The second time, I was told this exact case had already been settled once before, and that I was still applying a different rule than the one already agreed on. That was what made me look at what the alert was actually supposed to mean rather than what it was mechanically doing.

What shipped

The fix narrows the urgent category to fire only when the server that fails is one we actually operate. Everything else, including a customer’s own dead third-party host, still shows up in every scan, so nothing goes silently unreported, it just never pages as an emergency again. The change also added a way to mute a specific domain by hand with a reason attached, so a known, expected case like a pending migration can be acknowledged once instead of re-litigated on every page. The lesson wasn’t about the alert’s code being wrong in isolation. It was that restating a trigger condition as an answer, instead of checking the assumption underneath it, is exactly how a correct-sounding defense keeps a wrong alert alive.