The morning the second reminder email went to all 205 customer-managed domains that still had not moved over, I found out that the check behind our tracking screen had never looked at about a third of them.

A domain is the web address a customer’s site lives at, and those 205 were ones our customers manage themselves that still pointed at the old setup. A nightly check visits each one and records whether it has moved. The screen shows the results, and I had been reading that screen as the whole list, because nothing on it said otherwise.

What the screen was not saying

The nightly check has a one-hour limit. It was hitting the limit and stopping, with no error, no warning and no record that it had stopped early. Every site it reached had a row on the screen, and those rows looked exactly like proof of coverage. They were only proof that some work had happened before the clock ran out.

Shipping the reworded warnings, sending the email and finding this all fit inside one 39 minute session that ended at 8:56. I filed the finding along with three other follow-ups before closing it.

A blank means three things

A site with no problem showing could mean it was checked and fine. It could mean it was checked and still needs to move. Or it could mean the check never got to it. The screen could tell the first two apart. It could not tell the third from either.

For a campaign whose whole point is who still needs to move, that third case is the entire question.

The fix I did not pick

The obvious fix is a longer time limit. That only moves the line. A bigger limit still cuts the check off before the list is covered, and now it happens on a busier night where nobody is looking.

What the check has to record is three things: how many it expected to cover, how many it reached, and whether it finished. Only when all three agree does the list count as coverage. That day’s commit log also shows a change so the scan raises an alert when it gets killed, which is the silence that started this.

Where AI fits

Claude is quick at reading stored results and counting them against a list, and that is the question that exposes this kind of gap: not how many results exist, but whether every one of the 205 is accounted for by a run that finished. It only answers that if someone asks it to count expected against actual. Asked to summarise what the screen showed, it would have summarised a third of a list as if it were the list.

The human decision

Who receives a message is a decision about people, so a person decides whether the list behind it is complete. An assistant can draft the email and tidy the list. It cannot know the list was cut off unless the check says so.

The lesson

A check that a clock can stop has to say how far it got: expected 205, reached some number, finished or not. Until it says that, the rows it left behind are evidence that work happened, not evidence of coverage.

The Build Log companion covers the timeout, the missing run state and why extending the limit was the wrong repair.