---
title: "Making Firewall Rules From Live State"
canonical: https://dxdev.com/blog/2026-08-31_self-generating-firewall-rules/
datePublished: 2026-08-31
---
A firewall change log can cite the newest ruleset version and still be badly wrong, because two other summaries of the same ruleset sat beside it and the drift check never read either one. The live ruleset was v31 with 13 rules, the rule ID index in the same file still said v27 with 11 rules, and a second notes file said 8 rules at ruleset version 18 while teaching that the known-good-automation skip was phase-only, a claim that stopped being true in v25. The check compared one number, the newest `Ruleset version N` in the change log against the live version, so it reported four versions of drift and missed the docs that mattered, including a crawler block added at v28 that appeared in none of them.

The check came from a documentation backfill. We keep the firewall's custom ruleset in a markdown file with a change log at the bottom. The log had fallen behind, so I reconstructed versions 15 through 27 by diffing consecutive snapshots from the Cloudflare rulesets version API. I also refreshed the rule ID index in the same file, which still said 6 rules at v11 when the live ruleset had 11 rules at v27. Then I added a drift check to the weekly pre-release scan so the log could not quietly fall behind again.

The check watched the newest `Ruleset version N` citation in the change log and compared it to the live version.

## What the check could not see

Two other summaries sat beside the change log, and the check never read either of them. The rule ID index in the same file said v27, 11 rules, against a live v31, 13 rules. A second doc, the firewall notes kept alongside the agent instructions, said 8 rules, ruleset version 18. It also still described the known-good-automation skip as phase-only.

That last claim stopped being true in v25. It is the exact misunderstanding that took a payment webhook down earlier, so the stale doc was teaching the mechanism that had already caused an outage. A doc can cite the right version and still teach the wrong mechanism.

The docs also knew nothing about a crawler leak fix. The bypass rule skips on `cf.client.bot`, which is true of every verified crawler, so ClaudeBot and Amazonbot were waved past before the training-crawler policy was consulted. The fix was a new block rule placed above the skip, and it went in at v28. My change log stopped at v27. The skips are no longer contiguous, because rule 2 blocks AI training crawlers and must sit above the verified-bot skip.

## Backfilling v28 to v31 got the facts wrong

My backfill of the newer versions had errors of its own. I dated v28 and v29 to 08-27 and attributed v30 (a `.php` exclusion added to the skip rule) to the crawler ticket. The version API says v28 and v29 both landed 2026-08-25 and v30 landed 2026-08-26, and the crawler ticket's own branch never mentions v30 at all.

There was a second cost. The crawler ticket's branch had its own unmerged v28 and v29 entries, and merging it collided with my rewrite of the same file. That is the hazard the v29 entry itself warns about, two sessions editing this ruleset in the same hour, arriving inside the doc that warns about it. I kept my structure but took the branch's entries verbatim. They carried the measured numbers, 431 served and 0 blocked, 438 skip events, and the point that lowercasing the match would block 12,406 wanted Claude-SearchBot requests a day. My reconstruction had none of that. I then had to write a correction on the crawler ticket, and v30 now reads "no ticket record", with the nearest open question named as context and marked as a reading, not a record.

## Generating the index

So the index is generated. A local command builds it from the live ruleset and writes it between markers in the file. The command's `--check` mode asserts that the committed block still equals what the API returns. That single equality covers rule count, ids, order, actions, enabled state and skip scope together, so none of them needs its own check.

The prose needed a different treatment, because it cannot be generated. Counts in prose, in both docs, are now linted against the live ruleset. The lint applies only above the change log, since a change log records past states by design and a stale number there is correct. Any skip that is not `ruleset:current` gets reported on sight. The drift command also counts each disagreement as an anomaly instead of printing one advisory line.

I also rewrote the intent table in the firewall notes to the real 13 rules and relabelled it as intent, not inventory. A short section now says which of the three homes owns which fact. Nothing had said so before, which is why they disagreed.

## Verification

The command now reports "docs agree with the live ruleset (v31, 13 rules)". 1,661 tests pass, 7 of them new. They pin the failure modes from that day:

- a stale index while the change log is current
- the staff doc being checked at all
- the generated block not reporting its own count
- a phases-only skip being flagged

I verified the merge by reading both files back from origin, not from the local tree.

One question stays open. A second opinion argued the two docs should merge, since "private" means little once both sit in a staff-visible repo. That is a structural call in the middle of a knowledge-base migration, so I left it alone.

The check that said live 31 against doc 27 was correct, and the ruleset it never read was the one that mattered: 13 live rules where the index claimed 11, and a notes file that still taught the skip as phase-only after v25 had already changed it. The number that gave the drift away, four versions, was the smallest part of it, because the crawler block that landed at v28 above the verified-bot skip appeared in no doc at all. The index is now generated from the live ruleset and asserted equal to it, and the prose counts are linted against it, so a doc that still says 13 rules after the ruleset moves on fails the scan instead of passing on the strength of a current change log.
