A customer asked who set up their domain, and checking the answer showed our migration tracker was wrong about 85 sites. The old server address gets switched off in five days, and 96 domains still point at it.

That question sent me to the domains table of our legacy SaaS product. It has clientManaged, setupState and migrationEvidenceState columns, and the tracker is supposed to be the summary of them. Those columns had been telling us which customers were done. For 85 sites they were stale. So before anyone got a warning email, I built a dossier of 154 domain rows from live DNS, RDAP and WHOIS instead of the columns. The rest of this post is what that dossier did to my own assumptions.

The first wrong assumption: DNS is the answer

The lockdown rule is simple. The apex A record should point at the new target, and www should be a CNAME to the platform host. The old origin address is the one being retired. I recomputed a set of 33 domains whose apex, www_cname and www_a matched none of the three known targets and were not dead DNS. They resolved to website builders, a registrar’s parking page, a CDN, a big cloud host and assorted others.

I read that as “these people moved on.” I handed a research agent a read-only task built on that premise. Probe each domain over HTTPS, apex and www, follow redirects, and classify it as LEFT, PARKED_OR_PLACEHOLDER, STILL_OURS, UNRELATED or UNREACHABLE. The buckets were built around the assumption that a foreign A record means a foreign site. I expected LEFT to be the big one, and I planned to pull those rows out of the send list.

The probe output broke that. One basketball club’s apex resolved to a CDN, which looked like a clear “moved on.” The probe returned status 200, and the final URL was our own team page:

url=https://<club-domain>/ status=200
final=https://<platform>/teams/default.asp?u=<CLUB-USERNAME>&s=basketball
server=<cdn>

Several other domains looked the same. Their DNS pointed at a registrar’s forwarding service, and the registrar sent an HTTP redirect back to our platform. They served our site in full and matched none of our DNS targets. My exclusion list would have dropped live customers from the warning.

The fix was to stop classifying on DNS alone. Each domain got a probe, a redirect-chain check, and a second check on whether the account was alive. The account check fetches https://<platform>/?<username> and looks for a 302 to a team page. That is how a baseball club’s row could be marked NXDOMAIN with a WHOIS-confirmed “not registered” and still show a live account. The domain on file was a typo. The account was real. Those rows went to EXCLUDE_FROM_SEND with the reason written on the row.

The second wrong assumption: we already had the contacts

The dossier said ~126 customers to email. I assumed the contact fields on the account record held the right people. They didn’t. The sample row had "email": null, "Billingemail": null and "Phone": null. The first audience build came back with holes I had to fill by hand.

I rebuilt it from scratch, working from who actually administers each site. Along the way I found customers who had already left the platform, domains no longer registered, and sites already migrated. All of them came out of the list. What went out was a personalised warning to 94 accounts, covering 216 people, each naming that customer’s own domain.

The same day I nearly repeated the mistake on a different send. There were 26 recovery emails for customers whose expiration notices hard-bounced during the reverse-DNS outage. The drafts I’d prepared were identical for all 26. But those people had originally been sent three different versions, depending on how far along their renewal was. Account flags couldn’t tell me which version each person got. The real send log could, so I used it. It also caught one person the system had already re-contacted.

What the numbers looked like at the end

StageCount
Rows in the dossier154
Domains matching none of our three targets33
Planned recipients~126
Accounts actually warned94
People reached216
Sites the tracker had wrong85

The pattern shows in the shrinking counts. Every time I replaced a field I trusted with a fact I could probe, the audience moved. Some people came out of it, some went in, and a few moved to a different message. Twelve customers needed wording that didn’t fit any template, so they got their own ticket. Another ticket covers stale deadlines in the mailer copy.

We also shipped a production change from the same investigation. When domain forwarding is blocking the change a customer is being asked to make, the page now says so, names their registrar, and shows the same detail to support staff. That is the case that burned me in the probe run. It now surfaces where the customer and support can see it.

The tracker fix is filed. So is the plan to keep those 96 sites alive rather than let them go dark. The one thing I’d carry into the next migration is to treat every status column as a hypothesis, and to run the probe before the send.