The alert said 83 domains had migrated to Cloudflare. I checked the underlying dashboard query first, then checked reality, and the two didn’t agree.
The board that lied by omission
The migration console, an internal admin page with its own tracking ticket, renders a per-domain lifecycle: five dots, set up, live, secured, apex valid, old SSL removed. The www column pulls from Cloudflare’s custom-hostname API, so it shows active the moment CF issues a cert and accepts the hostname. That part is real. The bug was in what the board implied about the other half of the domain.
Every one of those 83 rows showed ✓ Cloudflare in the www column and a green dot string, which reads as “done.” But www on Cloudflare says nothing about the apex. The apex is deliberately kept off Cloudflare and served through an off-origin relay instead, on its own state (c.apexState), tracked separately in the ledger. The board had two independent trust domains, CF’s own state for www and our lifecycle ledger for apex, and it was letting the first one carry visual weight for both. A customer whose www went green looked migrated on the dashboard face, even when their apex CNAME had never been touched and was still resolving to us directly at origin.
I didn’t catch this by reading the ASP. I caught it by driving the live page and diffing what it rendered against DNS truth:
[...document.querySelectorAll('#cfListWrapper table tbody tr')].slice(0,3).map(r => { const tds = [...r.children]; return { domain: tds[2].innerText.trim(), state: tds[4].innerText.trim(), www: tds[7].innerText.trim() };})Then, for the same rows, resolving each domain against its own authoritative nameservers rather than trusting our cached zone data:
riverbendyouth.org ns03.domaincontrol.com CNAME riverbendyouth.org NOT-DONEpalmettobaseball.org ns22.domaincontrol.com CNAME palmettobaseball.org NOT-DONElakeshorecrusaders.com reseller-ns... A <origin> NOT-DONEAll 83 that the board marked as fully migrated were, at the apex, still pointed at us. Zero of the 83 were actually complete. I filed a ticket for the finding.
The wrong turn: trusting the cached zone data first
My first pass at fixing the false positives didn’t touch DNS resolution at all. I assumed the ledger itself had a stale write, some race where the apex-valid dot got set true without the CNAME check actually running, so I went and audited the lifecycle write path and the cron that updates apexState. That cost most of an hour and turned up nothing wrong, because the code was fine. The lifecycle state machine correctly required a live apex check before flipping the dot. The problem was that nobody had run a fresh check against authoritative nameservers in weeks. The cached rows in our DB were just old, sitting there as green dots for domains whose DNS had since drifted or never moved. I was debugging the write logic for a table that was accurately recording stale reads.
Once I gave up on the write path and just re-resolved live, the real picture showed up in minutes.
Measuring the other 134
With the false positives fixed, I ran the same live-DNS check against the full remaining list, 134 domains, not just the 83 that had been miscounted. Friday’s migration mailer had gone out to all of them. Live resolution showed 3 had actually acted on it. Three, out of 134, four days after the email.
Worse, a third of that list can’t execute the standard instructions at all. The runbook assumes the registrar exposes a normal DNS panel the customer can edit. For a chunk of these domains, the nameserver belongs to a reseller layer we don’t have a support path into, in one case a defunct one. lakeshorecrusaders.com sits under a white-label reseller storefront, and both of the reseller’s own support domains fail to resolve. The wholesale registrar one layer up still exists, but our mailer told the customer to go to a login page that no longer has anyone behind it. Another customer, on prairiebeltbaseball.com, was chasing the wrong domain entirely: her live account runs on prairiebelt.com, a different name on the same GoDaddy nameservers. She’d have gotten a login prompt and no way forward, and the mailer gave her no reason to suspect she had the wrong domain.
I filed three tickets. One for the reseller-blocked cohort, since it needs its own runbook branch rather than a retry of the same instructions. One for the wrong-domain case. And one for the underlying process gap: the wall detector that generates migration alerts had no mechanism to stop alerting on a domain once its deadline had already resolved one way or the other, so stale rows kept surfacing as active work days after they’d been handled or become unreachable through normal means.
What the dashboard actually needed
The fix wasn’t a smarter migrated count. It was making the board honest about which of its two states it was actually reporting. www green means Cloudflare accepted the hostname. It has never meant the apex moved. Once the UI stopped letting one column’s color imply the other column’s state, the count dropped from 83 false completions to zero, and the real number, 3 real actions out of 134 emails sent, was the one that mattered for deciding what to do next: not “send another reminder,” but “route the reseller-blocked third down a different path, and fix the wall detector so it stops reporting solved problems as open ones.”