---
title: "The Page Was Broken for Five Days. Our Monitors Did Not Catch It."
canonical: https://dxdev.com/blog/2026-09-26_broken-for-five-days-every-monitor-said-fine/
datePublished: 2026-09-26
---
One page in the admin was down for five days. Our monitors did not catch it.

The page lists sites. From September 11 to September 16 it threw a fatal ASP 0174 error for every site we tried. The server still answered with HTTP 200. I found it on the 16th.

## Why nothing flagged it

The daily uptime watcher counts Cloudflare edge 5xx minutes. A 200 with an error page inside it never adds to that count.

The weekly health report covers business numbers, not page health.

The application error log had no row for the crash. On the 16th the error log for the three days before it was checked. There were zero rows for that page, and it had been fatally broken the whole time.

The weekly site crawl only runs on Tuesdays, and it missed a page that rendered a fatal error inside a 200.

Every monitor we had keyed on a status code or a log row. This failure produced neither.

## What broke

A blank value reached a photo path and produced a double slash. `Server.MapPath` rejects that path. A commit on September 11 introduced it, and the fix has shipped.

## What we built instead of a new prober

The first plan was a separate prober for about 40 admin URLs. On September 26 we dropped it and extended the weekly crawl instead.

A page now fails if its text, or the background requests it makes, carries `ASP 0174`, `DEBUG MODE: Error`, an internal error code, or `Invalid object name`. It also fails if its main nav is missing. The crawl compares each run to the last one and alerts only on failures that are new.

After a prod push, a 13-page check runs 10 to 20 minutes later. It notices the push by watching the prod branches, so a slow build can mean it checked the old code. The full crawl stays weekly and is switched back on.

## What the crawl flagged

A partial crawl of the staff pages flagged four that render the internal error page inside a 200. The error log for the last 30 days was checked. Each page had three rows, all from the crawl's own visit, none from staff. They only error when opened with no parameters, so they are crawl artifacts.

## What I haven't seen yet

As of September 26, a real push had not yet triggered the post-push check, and the Discord alert had only been tested against a mock.
