At 8:56 AM, the migration console was telling a partial truth about 205 client-managed domains.
We had just shipped clearer migration warnings and sent the second-notice email to all 205 domains that still had not migrated. The console had a nightly scan behind it, so the obvious question was whether that list represented the full population. It did not. The scan had been dying at its one-hour limit, and roughly one-third of the domains had never actually been checked.
Nothing had raised an alarm. There was no obvious failure in the console. We had rows for domains the scan reached, and those rows looked like evidence of coverage. They were only evidence that some work had happened before the clock ran out.
That distinction changed the audit completely.
The one-hour boundary was part of the data
A migration verifier is not useful because it can inspect a domain. It is useful because it can make a reliable statement about a population. Once the job stops at sixty minutes without recording an incomplete run, it cannot make that statement.
The failure mode was quiet because the timeout lived below the console’s idea of success. The scheduler stopped the process at its time limit. The console showed the persisted results that existed. Neither component represented the missing part of the run as a first-class outcome.
That created an especially bad interface for an operational task. A blank result can mean several things:
- a domain was checked and needs no action,
- a domain was checked and needs migration,
- a domain was not checked because the job never reached it.
The console could distinguish the first two cases. It could not distinguish the third. For a migration campaign, that is not a missing detail. It is the central question.
Auditing the list, not the stored rows
The useful audit was not, “How many scan results do we have?” It was, “Can every one of the 205 domains be accounted for by a completed run?”
That is a different query, even when the storage is the same. Counting result rows tells you how much work produced data. Counting the expected audience against a terminal run state tells you whether the verifier covered its job.
The verifier needed an explicit contract closer to this:
expected_domain_count = 205processed_domain_count = count(results for this run)run_state = complete | incomplete | failed
coverage is valid only when:run_state == completeand processed_domain_count == expected_domain_countThe exact field names are not the point. The missing mechanism is. A verifier must preserve the execution outcome alongside its per-domain findings. If it does not, a time limit converts unprocessed work into invisible work.
That is why the audit showed false coverage. The console was not lying about the domains it displayed. We were asking it a question that its data model could not answer.
Why a longer timeout is not the repair
The tempting response is to extend the one-hour limit. That might reduce the frequency of the failure, but it does not change the semantics. A larger timeout still creates a boundary where the scan can stop before the population is covered. It just moves the boundary farther out.
Treating a timeout as an ordinary infrastructure failure is not enough either. This was not merely a job that needed attention. It was a job whose partial output stood behind a customer communication that had already gone out. The critical state was not only that an execution stopped. It was that the output was incomplete and could not support an all-clear claim.
The stronger design is to make incompleteness durable and visible. Keep the expected set for the run. Track the processed set. Record a terminal outcome that cannot be mistaken for success. Then resume or rerun only from a state that preserves the gap instead of overwriting it with a fresh-looking partial result.
The second-notice email made the stakes clear. We were not building a dashboard ornament. We had already sent that email to every domain still unmigrated, and for roughly one-third of that audience the verifier behind the console had never actually checked the domain.
The scan did not merely time out. It stopped being a scan.