One page in the admin was down for five days. Our monitors did not catch it.
The page lists sites. From September 11 to September 16 it threw a fatal ASP 0174 error for every site we tried. The server still answered with HTTP 200. I found it on the 16th.
Why nothing flagged it
The daily uptime watcher counts Cloudflare edge 5xx minutes. A 200 with an error page inside it never adds to that count.
The weekly health report covers business numbers, not page health.
The application error log had no row for the crash. On the 16th the error log for the three days before it was checked. There were zero rows for that page, and it had been fatally broken the whole time.
The weekly site crawl only runs on Tuesdays, and it missed a page that rendered a fatal error inside a 200.
Every monitor we had keyed on a status code or a log row. This failure produced neither.
What broke
A blank value reached a photo path and produced a double slash. Server.MapPath rejects that path. A commit on September 11 introduced it, and the fix has shipped.
What we built instead of a new prober
The first plan was a separate prober for about 40 admin URLs. On September 26 we dropped it and extended the weekly crawl instead.
A page now fails if its text, or the background requests it makes, carries ASP 0174, DEBUG MODE: Error, an internal error code, or Invalid object name. It also fails if its main nav is missing. The crawl compares each run to the last one and alerts only on failures that are new.
After a prod push, a 13-page check runs 10 to 20 minutes later. It notices the push by watching the prod branches, so a slow build can mean it checked the old code. The full crawl stays weekly and is switched back on.
What the crawl flagged
A partial crawl of the staff pages flagged four that render the internal error page inside a 200. The error log for the last 30 days was checked. Each page had three rows, all from the crawl’s own visit, none from staff. They only error when opened with no parameters, so they are crawl artifacts.
What I haven’t seen yet
As of September 26, a real push had not yet triggered the post-push check, and the Discord alert had only been tested against a mock.
AI Skills
Use this lesson with the AI assistant you already use
An admin page rendered a fatal error inside an otherwise normal, successful-looking response for five days. The uptime check watched for server error codes, the error log never received a row, and the weekly crawl missed it, so none of them caught a failure that produced none of the signals they were watching for.
Paste the prompt, share only the context needed to answer it, and treat the result as a draft for your review. Do not include confidential information or let an AI assistant make changes without your approval.
Optional: for a visual report and saved memory, run /dxdev first.
Don’t have it? Get it at dxdev.com/skills/dxdev. The prompt works without it.
dxdev LESSON · paste into your AI coding agent
LESSON: A Monitor Only Catches the Failure Shape It Was Built to Recognize
SOURCE: dxdev.com/blog/2026-09-26_broken-for-five-days-every-monitor-said-fine
WHAT HAPPENED: A page crashed while still returning a normal, successful response status, and stayed broken for five days before a person noticed it. The uptime watcher only counted true server errors, the application's own error log received zero rows for the failure, and the site's existing weekly crawl missed it. The fix extended the existing crawl to inspect page content and background requests for known failure signatures, check that expected page structure is present, run a reduced check after every production deploy in addition to the full weekly pass, and alert only on failures that are new since the previous run so already-known issues do not create ongoing noise.
THE RULE: Before trusting a monitor's "all clear," identify the specific failure signals it actually watches (a status code, a log row, a page load) and confirm those signals are the ones the failures you actually care about will produce; a monitor that has never missed anything may simply never have been tested against the failure shape it cannot see.
CHECK MY CODE, then report PASS or FAIL with file:line for each:
1. Any monitoring or health check whose pass condition is a status code or "page loaded" rather than an inspection of the actual returned content.
2. Any error logging path that could be bypassed by a crash occurring before the logging code executes, leaving a real failure with zero log evidence.
3. Any automated check that would re-alert on every run for a known, already-tracked issue rather than diffing against a previous baseline and surfacing only new failures.
THEN PRINT: a table (check, PASS/FAIL, evidence, fix) + a verdict (applies / partially / OUT_OF_SCOPE / no) + the single most important next action.