One domain sat in pending_validation for about thirteen hours, and the validation error field was null the entire time.
The field that should explain a stall was empty. I was moving customer domains onto Cloudflare and our own relay, one batch at a time, without taking any site down. For each domain the routine is the same. Register the hostname with Cloudflare, publish an _acme-challenge TXT record so the certificate authority can prove the domain is ours, wait for the certificate to go active, and only then flip the traffic. The site stays on its old origin until the certificate is live, so a domain that sticks is never an outage. It just does not move.
What I checked, and what I concluded
The stuck domain had my TXT record published, and it resolved from outside. Our DNS looked clean and Cloudflare said nothing was wrong. When I rechecked at around the twelve hour mark, Cloudflare was listing two validation tokens for the hostname, and I wrote that down as token rotation. In the same note I still called it a Cloudflare or certificate authority side wedge, because our DNS was clean.
That conclusion was wrong. Cloudflare rotates the expected token if the first validation window lapses, whether from a slow publish, DNS propagation, or its own backoff. My TXT record was the token from before the rotation. It existed and it resolved, and it was not the record Cloudflare was now waiting for. What I had checked was whether my record was published. What I needed to check was whether it matched.
The ramp-up batch made it a pattern
Later that day I ramped to 25 live domains in one batch. All 25 went through with zero downtime, but four of them stalled in the same state, pending_validation with a null error, and it was the same cause. In that batch, 16% had their token rotated out from under them.
For those four, I asked Cloudflare for each hostname’s current expected token, published it alongside the stale one, and waited again. Cloudflare accepts any matching TXT record at that name, so I did not have to delete the old ones. All four validated and flipped.
The thirteen hour domain was different. Publishing the current token did not move a hostname that stale, so I deleted it and registered it again from scratch. The fresh registration got a new validation window and the certificate was active in about a minute.
Turning the check into a command
Four more stalls in one batch made it a pattern, so the check became a first class step. A reconcile command returns Cloudflare’s current expected token for any domain. The wait step now prints that token for anything still pending past its window, instead of timing out and reporting nothing. Anyone can compare it with what sits in the zone and see the mismatch in one command.
I also started recording how long each domain takes to go active, so a batch can be planned against real numbers.
The empty error field was where I stopped looking that day. The step that catches it now shows the token Cloudflare wants next to the one I published.