The blank error pages had been going on for weeks before anyone tied a number to them: 39 error-minutes a day, averaged over a 7-day window, out of 41 million requests. That ratio is why it hid so long. 40,000 bad requests out of 41 million rounds to nothing in a dashboard. But those bad requests weren’t spread evenly, they landed in clusters, and if you were one of the people browsing during one of those clusters, the site was just down.
A customer’s “site is down” email is what forced the real hunt. Four passes ran in parallel: log analysis on the origin, live traffic sampling, a Cloudflare analytics pull, and a counter watch deployed straight onto the box. The origin came back clean every time. Zero 5xx from the app, p95 latency actually dropped during the biggest burst, no app-pool recycles, no crashes. Whatever was breaking wasn’t the server.
The wrong turn: rate limiting, fast
Once the shape looked like “floods overwhelming a shared front door,” the obvious fix was a Cloudflare rate-limiting rule, block or challenge any IP over some threshold. It’s a Pro-plan feature, it would have stopped every flood we’d observed, and it’s a five-minute change. I nearly shipped it on that logic alone.
Then I remembered the WAF already had a track record of doing this kind of damage: a prior rule had silently 403’d server-to-server webhook POSTs, empty-user-agent traffic that all looked like bots because it never carried a browser signature: a payment processor’s asynchronous callback webhooks. That rule sat live for a while before anyone noticed payment confirmations were failing, because a blocked webhook doesn’t throw an error a human sees, it just quietly never arrives.
So instead of shipping the rate limit, I ran it against real traffic first. Two studies, read-only, no config touched: one baselining what legitimate per-IP burst rates actually look like (the platform’s audience is schools and leagues, so 200 parents on one gym wifi hitting refresh during a tournament is a real, legitimate burst, not an attack), and one checking known server-to-server callers, payment callbacks, calendar feeds, importers, for the low-UA-diversity, fixed-IP, bursty pattern that a naive threshold would flag as abuse. If I’d shipped the rule on the first instinct, it would have broken exactly the traffic the WAF had already broken once.
What the flood actually was
The counter watch pinned it: during a 6,948-request burst, 74% came from a single IP hammering /errors/Connection.asp in a redirect loop, at up to 353 req/sec, layered on top of a WordPress REST scan sweeping dozens of vanity domains at exactly 72 hits each. Cloudflare passed all of it straight to origin. About 6% of the surge got reset (520) or timed out (522) at the connection layer before it ever reached IIS, invisible to the W3C logs because a dropped TCP handshake never gets logged on the app side. That’s why the origin looked perfectly healthy the whole time: the failures were happening upstream of it.
The fix that shipped was staged, not a single rule: WAF skips for an SMS provider’s and a payment processor’s callback paths first, then .php path blocking (flat block, not a challenge, because challenging 7.5 million requests a week would add load at the exact bottleneck being relieved), then a per-IP rate limit on /errors/Connection.asp. A daily error-minutes metric went in alongside it, watching distinct minutes per day with at least one edge 5xx on a main-site host, not raw request counts, because request counts are exactly what hid this for weeks. Baseline 39.4/day, threshold alert at 15, paging only on state transitions so a quiet week stays silent and a standing problem doesn’t re-page every morning.
Three sources, one shared ceiling
The edge rules cut error-minutes by 92% almost immediately. Then they stalled, and 670 failed requests a day kept showing up. That’s the part that mattered: it wasn’t one bug with one fix, it was three independent things all filling the same finite burst capacity on the web box, which has no headroom to absorb concurrent spikes regardless of source.
Two of the three showed up in the leftover 670: 76% of it was two identifiable scanners, one flagged by user agent (l9scan/2.0, all from a single ASN, 35,971 requests since the rules shipped) and unmatched probe paths like /getcmd and /aaa9 that generated zero firewall events because nothing in the ruleset described them. The third source wasn’t a scanner at all: a hand-run SQL sweep against Query Store, and separately, a calendar-feed retry storm hammering .ics endpoints when an upstream skip rule missed ruleset: current scoping and let a legitimate feed get 403’d into retrying itself into a flood.
Whoever fixed the scanner rules first saw their error rate drop. Whoever was hitting the calendar retry storm or riding behind the SQL sweep saw nothing improve, because their burst was never the scanner’s burst. Same symptom, same dashboard metric, three unrelated causes sharing one ceiling. Six follow-up tickets came out of that pass, one per source, because the fix that helps most customers doesn’t help the ones behind a different queue.