The Watchdog Was Watching the Wrong House
The plan was to move 25 domains, including the main one customers use every day, off an aging server and onto new infrastructure with nobody’s mail or site going down in the middle of it. No dedicated operations team, just a careful inventory of every domain’s pieces: where it points, what mail routes through it, what has to move together so nothing breaks mid-swap.
The move itself went fine. Roughly six and a half hours, twenty-five domains relocated, no outage, no lost mail. That’s the part that looked, from the outside, like the whole story.
The part that actually mattered came from checking whether the safety net under the move was real. Before flipping anything, I looked at what the existing health monitor, the thing that’s supposed to tell someone the moment a site goes down, was actually watching. It wasn’t watching the business’s live domains. It was watching a personal address and a couple of test domains, some of them expired, that had nothing to do with what customers were using.
That’s not a bug the migration caused. That monitor had been pointed at the wrong things for a while, quietly, with nobody noticing, because a monitor that’s running looks the same from the outside whether it’s watching the right thing or the wrong thing. Green means green. It doesn’t say green for what.
Here’s the uncomfortable part: if the migration had gone badly, if one of those twenty-five domains had actually gone down mid-move, the monitor watching the business would have said nothing, because it wasn’t looking at that domain in the first place. The safety net people believed existed didn’t cover the thing it needed to cover. That gap could have sat there indefinitely, waiting for a bad day to expose it, and this migration only found it because someone finally went and read what the monitor’s configuration actually said instead of trusting that it existed.
The fix was sequencing, not just repointing. The monitor’s cutover to the real, customer-facing domains happened only after the new routing was already verified to work. That order mattered, because a green check on the old, wrong targets during the actual move could have been mistaken for real coverage at the exact moment coverage mattered most. Getting the order backward would have meant trusting a working-looking light for a system that wasn’t actually being watched yet.
If you have a monitor you trust
Having a monitor is not the same claim as the monitor watching the right thing. It’s worth a five-minute check: open whatever is doing your health checking today and read, literally, what addresses or systems it’s pointed at. Compare that list against what your customers actually use. If those two lists don’t match exactly, you have a gap that behaves exactly like coverage until the day you need it.