---
title: "When Your Diagnostic Says the Customer Broke It"
canonical: https://dxdev.com/blog/2026-09-07_diagnostic-ui-false-blame/
datePublished: 2026-08-27
---
The alert wasn't a bug report. It was our support lead, on the phone with a customer for hours, because our own screen was telling her he was wrong.

One customer had pointed his DNS at us correctly, and his site was already fully live and secure. But the staff panel we use to check domain setup showed a hard red "Not on Cloudflare yet," with a note that he still needed to update his DNS. Right underneath that red banner, on the same screen, sat the actual DNS lookup, live and current, showing both records already pointed at us. The panel was contradicting itself and picking the wrong side.

The cause was one function trying to answer two different questions with one yes-or-no flag. Part of our system makes a live web request to the customer's domain to confirm we're actually serving traffic there. That request can fail for reasons that have nothing to do with the customer: our probe times out, a certificate can't be validated in that moment, the connection gets refused. Whatever the reason, a failed check and "the customer's DNS is wrong" got written to the exact same true/false value. So when our own network hiccuped, the screen didn't say "we couldn't confirm this." It said the customer hadn't done his job.

I went looking for a root cause and found a red herring instead. My first theory was that the security certificate chain on the customer's domain used a newer root certificate our server might not trust, something that would fail every time and need a real fix. I ran the exact same check by hand from our own server against his domain. It came back 200 OK, a clean success. The certificate root I'd suspected was sitting right there in the server's trusted list. My theory was wrong, and worse, it would have sent me down a day of chasing a certificate problem that didn't exist. The actual failure was a single blip, a one-off hiccup with no retry behind it, and that blip is what got permanently stamped onto a customer's account as "his DNS is broken."

That's the part worth sitting with if you run a business on top of software you didn't write yourself. A diagnostic screen that fails silently doesn't fail as "unknown." It fails as an accusation. The screen doesn't have a way to say "I don't know." It only has the options it was built with, and if nobody built in a way to say "our check didn't work," the tool will confidently blame someone. In this case it blamed a paying customer, out loud, to a staff member who then spent hours defending our system to him on the phone for a mistake he hadn't made.

The fix wasn't just "retry the check." We did add one retry, but a retry only makes the blip less likely, it doesn't stop it from lying when it happens again. The real fix was giving the check a third answer. Instead of only "yes, it's working" or "no, it's not," it can now also say "I couldn't tell." That third state gets treated completely differently on screen: a real DNS problem still shows red and still says the customer needs to fix something. But "we couldn't confirm it" now shows as our open item, not the customer's, and it explains that the setup often finishes on its own within a few hours, since these DNS changes need time to spread across the internet anyway.

We also split the panel so it can't contradict itself. The question "did the customer point his DNS correctly" now reads directly from the DNS record. The question "are we actually serving the site yet" is tracked and shown separately. Before, one broken probe could paint over a correct, live DNS record with a red banner nobody could explain. Now the two facts can't overwrite each other, and whoever's on the phone can read out which half is unproven instead of guessing.

We tested it against seven different real situations pulled from actual accounts, including the one where a customer genuinely hasn't pointed his DNS yet, to make sure that case still shows red and still says so. It does. What changed is only the case where our system doesn't actually know.

The bigger habit this forced was checking the machine before trusting my own explanation for it. I had a plausible, specific story for why the check failed, and it was wrong. If I'd shipped a fix for that story instead of testing directly against the live server, we would have burned time patching a certificate issue that was never there, while the panel kept quietly blaming customers for a problem on our end.
