At 12:18 AM, we had just moved the apex and www records for a guided customer home onto Cloudflare Pages, with mail still intact, when I noticed the awkward part of the architecture: our status page depended on the same edge network as the thing it was meant to explain.

That is not redundancy. It is a second tab in the same outage.

The public site was deployed, the DNS cutover was done, and the mail records had survived the change. That was a useful result, but it made the dependency graph painfully obvious. If Cloudflare had a broad incident, visitors could lose the site and we could lose the normal path for telling them what had happened. A status page behind the same CDN, served by the same provider, and reached through the same DNS path is available only when the primary path is available.

I did not find this through a synthetic monitor or an incident review. I found it while looking at what we had just shipped. The main deployment was healthy. That was exactly why the problem was easy to ignore.

The status page that could not speak

The first version of the plan was obvious: keep a /status route close to the product, publish updates there, and link to it from the footer. It looked tidy. It also failed the only test that mattered.

A visitor reaches that route through the same chain as the rest of the site:

browser -> DNS -> CDN edge -> static deployment

When the failure is inside the CDN or its shared control plane, the route has no independent component in that chain. Changing the copy, giving it a different path, or making it visually simpler does not change that. The page is still downstream of the outage.

We also considered putting the same status application on a different hostname while leaving it on the same provider. That lost for the same reason. A different hostname can help with a bad deploy, an application-route mistake, or a cache rule that targets one site. It does nothing when the common provider is unavailable.

The other discarded option was a third-party status product as the only escape hatch. That would have separated the infrastructure, but it would also have made an emergency update dependent on a vendor workflow and its account boundary. We wanted an independent publishing path we could understand end to end, not a dashboard we hoped to remember how to use while the primary system was failing.

Build the escape path while nothing is broken

The answer was a break-glass board on separate infrastructure. It is intentionally small. Its job is not to recreate the main site or give people a polished support experience. Its job is to publish a timestamp, an incident state, a short statement of what we know, and the next time we expect to update.

The important part is where it lives. The board is not deployed through the main Cloudflare Pages path, and its normal public address is not attached to the main site’s routing. That changes the chain to this:

browser -> independent hosting path -> status board

A broken application deployment can take down the product without taking down the board. A bad rule on the main edge can take down the product without taking down the board. A Cloudflare incident can take down the product without deciding whether the board exists.

I added one more path because DNS itself can be part of the failure. The break-glass board has a raw-IP fallback. If the normal status hostname cannot resolve or cannot be trusted during a provider incident, we can publish and circulate a direct IP address for the board.

That fallback is deliberately ugly. It does not try to preserve the normal brand, navigation, or application session. It is there to answer four questions: is there an incident, what is affected, what have we done, and when will we post again. In a real outage, that is the contract.

There is a tradeoff worth naming. Browsers expect a hostname for normal TLS certificate validation. A raw-IP fallback is therefore not a consumer-grade destination, and it should never be the first link we give people on an ordinary day. It is a documented emergency route for the case where the ordinary name is part of the incident. We keep the normal independent hostname as the primary status address and reserve the direct address for the narrow failure mode it solves.

Test the dependency, not the page

The diagnostic question changed after that. I stopped asking, “Does the status page load?” Every status page loads when the internet is working. The useful question is, “Which component has to stay healthy before we can publish a sentence?”

For the primary site, the answer includes the Cloudflare edge and the DNS configuration we had just changed for the apex and www records. For the break-glass board, the answer must not include either dependency. It needs separate hosting, a separate operational path for updates, and the raw-IP route documented before anyone is tired, annoyed, or trying to answer customers.

That last point is why I built this as part of normal deployment work. An outage is the worst time to discover that your status page is only a marketing page with a different color scheme. It is also the worst time to decide who can publish an update, where the address lives, or whether the fallback has actually been written down.

The main site can have every performance and security improvement we shipped in this pass. It can have a 363 ms LCP, a 1,253 ms prior baseline, CSP, HSTS, a sitemap, and a polished public surface. None of that helps if the provider behind it is unavailable and our incident channel is coupled to the same failure.

A status page is not part of the product’s design system. It is part of its failure system. I want ours to be reachable precisely when the rest of the site is not.