---
title: "When 83 Domains Look Done But None Were Complete"
canonical: https://dxdev.com/blog/2026-09-07_partial-migration-confidence-trap/
datePublished: 2026-08-03
---
The alert said 83 domains had migrated to Cloudflare. I checked the underlying dashboard query first, then checked reality, and the two didn't agree.

## The board that lied by omission

The migration console, an internal admin page with its own tracking ticket, renders a per-domain lifecycle: five dots, set up, live, secured, apex valid, old SSL removed. The `www` column pulls from Cloudflare's custom-hostname API, so it shows `active` the moment CF issues a cert and accepts the hostname. That part is real. The bug was in what the board implied about the other half of the domain.

Every one of those 83 rows showed `✓ Cloudflare` in the `www` column and a green dot string, which reads as "done." But `www` on Cloudflare says nothing about the apex. The apex is deliberately kept off Cloudflare and served through an off-origin relay instead, on its own state (`c.apexState`), tracked separately in the ledger. The board had two independent trust domains, CF's own state for `www` and our lifecycle ledger for `apex`, and it was letting the first one carry visual weight for both. A customer whose `www` went green looked migrated on the dashboard face, even when their apex CNAME had never been touched and was still resolving to us directly at origin.

I didn't catch this by reading the ASP. I caught it by driving the live page and diffing what it rendered against DNS truth:

```js
[...document.querySelectorAll('#cfListWrapper table tbody tr')].slice(0,3).map(r => {
  const tds = [...r.children];
  return { domain: tds[2].innerText.trim(), state: tds[4].innerText.trim(), www: tds[7].innerText.trim() };
})
```

Then, for the same rows, resolving each domain against its own authoritative nameservers rather than trusting our cached zone data:

```
riverbendyouth.org        ns03.domaincontrol.com   CNAME riverbendyouth.org        NOT-DONE
palmettobaseball.org      ns22.domaincontrol.com   CNAME palmettobaseball.org      NOT-DONE
lakeshorecrusaders.com    reseller-ns...           A <origin>                      NOT-DONE
```

All 83 that the board marked as fully migrated were, at the apex, still pointed at us. Zero of the 83 were actually complete. I filed a ticket for the finding.

## The wrong turn: trusting the cached zone data first

My first pass at fixing the false positives didn't touch DNS resolution at all. I assumed the ledger itself had a stale write, some race where the apex-valid dot got set true without the CNAME check actually running, so I went and audited the lifecycle write path and the cron that updates `apexState`. That cost most of an hour and turned up nothing wrong, because the code was fine. The lifecycle state machine correctly required a live apex check before flipping the dot. The problem was that nobody had run a fresh check against authoritative nameservers in weeks. The cached rows in our DB were just old, sitting there as green dots for domains whose DNS had since drifted or never moved. I was debugging the write logic for a table that was accurately recording stale reads.

Once I gave up on the write path and just re-resolved live, the real picture showed up in minutes.

## Measuring the other 134

With the false positives fixed, I ran the same live-DNS check against the full remaining list, 134 domains, not just the 83 that had been miscounted. Friday's migration mailer had gone out to all of them. Live resolution showed 3 had actually acted on it. Three, out of 134, four days after the email.

Worse, a third of that list can't execute the standard instructions at all. The runbook assumes the registrar exposes a normal DNS panel the customer can edit. For a chunk of these domains, the nameserver belongs to a reseller layer we don't have a support path into, in one case a defunct one. `lakeshorecrusaders.com` sits under a white-label reseller storefront, and both of the reseller's own support domains fail to resolve. The wholesale registrar one layer up still exists, but our mailer told the customer to go to a login page that no longer has anyone behind it. Another customer, on `prairiebeltbaseball.com`, was chasing the wrong domain entirely: her live account runs on `prairiebelt.com`, a different name on the same GoDaddy nameservers. She'd have gotten a login prompt and no way forward, and the mailer gave her no reason to suspect she had the wrong domain.

I filed three tickets. One for the reseller-blocked cohort, since it needs its own runbook branch rather than a retry of the same instructions. One for the wrong-domain case. And one for the underlying process gap: the wall detector that generates migration alerts had no mechanism to stop alerting on a domain once its deadline had already resolved one way or the other, so stale rows kept surfacing as active work days after they'd been handled or become unreachable through normal means.

## What the dashboard actually needed

The fix wasn't a smarter migrated count. It was making the board honest about which of its two states it was actually reporting. `www` green means Cloudflare accepted the hostname. It has never meant the apex moved. Once the UI stopped letting one column's color imply the other column's state, the count dropped from 83 false completions to zero, and the real number, 3 real actions out of 134 emails sent, was the one that mattered for deciding what to do next: not "send another reminder," but "route the reseller-blocked third down a different path, and fix the wall detector so it stops reporting solved problems as open ones."
