---
title: "I Fixed the Bot Floods, Then My Fix Blocked Real Customers"
canonical: https://dxdev.com/blog/2026-08-21_firewall-rules-false-positive-guard/
datePublished: 2026-08-21
---
Four edge rules took main-site downtime from about 39 minutes a day to zero within the hour. Later the same day I found out they were turning away real customer calendar feeds and search engines.

## Finding the flood

The symptom was chronic blank error pages. I root-caused it by live-sampling 12 bursts and watching the origin's on-box counters. They didn't move in any of the 12. The origin was verified healthy, so scanner-bot floods were overwhelming the shared network front door upstream of it. New visitor connections were getting dropped at Cloudflare, and the 520 error flood was the visible result.

The fix was four Cloudflare edge rules:

- A skip rule so payment and SMS callbacks never hit the other rules.
- A block on `.php` and archive extensions. I extended the existing archive-extension rule and made the match case-insensitive, after finding the original was case-sensitive.
- A block on scanner recon paths: environment files, git, WordPress admin and similar.
- A port block.

I validated the block rules against a week of production traffic before they went live. I had first staged `.php` as a challenge. Then I changed it to a flat block, because challenging 7.5M requests a week would add load at the exact connection bottleneck I was trying to relieve.

## The measurement problem

Downtime dropped to zero, but the old uptime probe could never have caught a failure rate as thin as the one I now needed to watch for. So I built two things: a daily watch with alerting, and a guard that checks whether our own rules are blocking real visitors.

## The allow rule that allowed nothing

When I checked the release, one allow rule turned out to be holding back only Cloudflare's own managed filters. It did nothing about my new custom block rules. Real customer calendar feeds and search engines were getting turned away.

The ticket had closed to staging with downtime at zero. That number was true of the scanners and of the real traffic I was accidentally cutting off, because a rule that blocks the wrong requests also produces zero error pages.

I fixed it, then confirmed the feeds were serving again while the scanner traffic was still blocked. I also filed a ticket for the WAF reference doc we were missing.

On Aug 26 one real customer calendar feed turned out to be wrongly blocked, and I fixed that too.

## Making the guard honest

The guard checks blocked requests. Those carrying a referer from a customer's `wp-admin` page could be real users being blocked. When I looked at them, they came from hosting IPs with a `WordPress/6.4.3` user agent, probing paths that don't exist on our side. That is scanners faking a referer, not real users. The same check gave the same answer on Sep 1 and again on Sep 28.

I also had to decide what counted in the metric. I corrected the measurements so comparisons stayed honest, and I had to decide explicitly whether one media host belonged in the metric before running the final scan.

## What it looks like now

As of 2026-08-27, main-site error-minutes were down 83% (6.4/day against a 37.7 baseline) over 44.6 hours. By Sep 28, five weeks on, I checked live:

- Main-site error-minutes for Sep 21 to 27 were 23 in total, about 3.3 a day, down 91% from 37.7 a day. Yesterday had zero.
- Errors across all customer sites were 8 yesterday, against a pre-fix average of 6,459 a day.
- All four rules were still in place.

The errors that remain are slow-origin timeouts, not attacks, and they belong to a separate ticket. I closed the flood ticket on that basis.

The original goal was fewer error pages, and the rules hit it within the hour. Zero error pages was also what rules that rejected real customers would have produced, so the count alone could not tell the two apart. A block rule needs two checks: the bad traffic is blocked, and the legitimate traffic is still served.
