At 4:10 PM, a migrated site had the right DNS records and no HTTPS.
The customer had completed the Cloudflare domain migration, but the custom-hostname certificate had stopped moving. From the outside, it looked exactly like ordinary DNS propagation. The domain pointed where it should. The staff interface did not surface the actual certificate condition. The site just could not serve HTTPS.
That distinction matters. “Wait a little longer” is acceptable advice when the thing still converging is DNS. It is bad advice when the certificate workflow is wedged. The customer sees the same broken site either way, but the repair path is completely different.
I fixed the affected certificate first. Then we turned the incident into a system that can identify the same state, repair it, and show staff what is actually happening.
The symptom was not the state
Our first input was blunt: a customer had done a Cloudflare migration and it had not gone through yet. That was plausible wording, because a migration has several moving parts. The obvious first suspect was a bad or incomplete cutover.
The audit changed the diagnosis. We listed all 1,282 custom hostnames and classified the failing domain instead of treating every moved hostname as a candidate for a retry. The domain was pointed correctly. Its certificate was the part that was stuck.
That gave us a cleaner model of the failure. A domain with no HTTPS after migration can look like DNS propagation, but a pointed domain can still be waiting on certificate issuance. The absence of a surfaced certificate condition does not mean there is nothing to do. It means staff need certificate state, not a generic migration state.
The incident was not a case for more monitoring. It was a case for better state separation. DNS state and certificate state had been collapsed into one vague idea of a migration being “in progress.” Once we could see the certificate state directly, the repair was straightforward.
We did not bulk retry the fleet
The tempting fix was to re-trigger every migrated hostname. I nearly did that during the audit. It would have made the immediate problem disappear from the queue, and it would have been reckless: bulk-reissuing across 1,282 custom hostnames would have caught domains still in the middle of a legitimate, still-converging cutover and forced them into a fresh issuance cycle they did not need, on the strength of one customer’s report. The instinct to clear the whole list at once came from wanting the incident closed, not from evidence that the whole list had the same problem. I stopped myself before running it, but the honest version of this story is that I got as far as considering it the fast path before the classification step talked me out of it.
The useful condition was not “hostname has moved.” It was stuck but pointed. That DNS gate mattered. Reissuing every moved hostname would mix up domains still completing a legitimate cutover with the narrow class that had a certificate failure.
So we kept the classification step in the design. The worker lists custom hostnames, identifies the stuck-but-pointed records, and re-triggers only those. That is the difference between an automated repair and an automated source of noise.
The other rejected option was to leave this as a staff-only procedure. We could have documented the recovery and told staff to open the provider console whenever a customer reported missing HTTPS. That would still leave a customer waiting for an issue that a scheduled worker could recognize in minutes. It also would not address the staff interface lying by omission.
We shipped both paths because they solve different parts of the same problem.
The recovery loop runs every three minutes
The production certificate worker now runs every 3 minutes. Its recovery step, Invoke-HealStuckCerts, does three things in sequence:
- Lists the custom hostnames.
- Filters for the stuck-but-pointed certificate condition.
- Re-triggers certificate issuance for that narrow set.
There is a cooldown ledger in the loop. Without it, a certificate that stays stuck across multiple scans could be re-triggered every three minutes. That is not recovery. It is an accidental retry storm with a nicer name.
We found a real implementation trap immediately after deployment. The cooldown ledger was named cf-heal-state.json and written into the worker’s queue directory. The queue scanner globbed *.json, so it picked up its own ledger and processed it as a bogus command.
The fix was small and specific: cooldown ledgers use .txt, not .json, in that directory. But the lesson is bigger than the suffix. A queue directory has a schema even when the filesystem is the schema. If the worker defines work as “every JSON file here,” then any JSON file placed there becomes work. State files need a namespace the consumer will not mistake for input.
We caught that in the live production log, fixed it, and verified the recovery loop end to end. This is why I prefer a repair that exercises itself immediately over a clean-looking implementation that waits weeks for the next incident to reveal its edge cases.
The staff interface got a real state and an escape hatch
Automation is not a substitute for observability. A worker can fix most cases in three minutes and still leave staff blind during the exception it cannot classify.
One release changed the domain dialog in two ways. It now shows the true SSL state, including a stuck certificate, and it gives staff a Re-trigger action. The button is not the primary mechanism. The scheduled recovery loop is. The button is the manual override for the moments when staff are already looking at the domain and need to act now.
That pairing matters. A raw provider status is useful only if the person reading it knows what to do next. A button without the state behind it is a ritual. We put the diagnosis and the recovery action next to each other.
The complete flow is now short enough to reason about:
migration completes -> custom hostname is pointed -> certificate state is checked -> stuck-but-pointed records enter the recovery path -> worker re-triggers issuance every 3 minutes, with cooldown -> staff can see the same state and re-trigger once when neededWe also documented the certificate flow and the surrounding infrastructure so this does not survive only as one person’s memory of a customer incident.
The original site was fixed within the same pass that found it. The worker that now runs every 3 minutes almost undid itself on day one, glitching on its own cooldown ledger because the queue scanner matched any *.json file in its directory, including the one the worker had just written to remember not to retry too soon. We caught that in the live production log and renamed the ledger’s extension to .txt. A self-healing loop that cannot tell its own memory apart from a customer’s stuck certificate is not healing anything; it took one production run to prove that distinction had to be enforced by a file suffix, not assumed.