---
title: "The Cloudflare audit found eight ip-handling bypasses. There was one I had to leave broken."
canonical: https://dxdev.com/blog/2026-05-27_cloudflare-audit-cant-fix-one/
datePublished: 2026-05-27
---
I was about to flip the platform's nameservers to Cloudflare. Phase 1b had landed days earlier in `www/library/src/Session-s.asp` and I had been treating that as the IP-handling fix. The session abstraction was reading `CF-Connecting-IP` first and falling back to `REMOTE_ADDR`. Every place that used the session for IP was now correct. I had tested the fallback on a non-CF environment, watched the logs come back honest on the live customer sites running CF, and put the nameserver flip on the calendar.

The thing I almost didn't do was a final grep.

It was a five-second thought. Phase 1b fixed the abstraction, but the abstraction is only protective if every site reads through it. I had not actually verified that. I ran a grep for `REMOTE_ADDR` across `www/` expecting maybe two or three legacy reads in dead corners.

Eight files came back.

## the audit

Six of the eight were live code paths. They read `REMOTE_ADDR` directly, never went through the session, and would have started logging Cloudflare's edge IP the moment the nameservers flipped. The kinds of things they were doing were the kinds of things that quietly stop working when the IP they see becomes wrong: a staff block list, a diagnostic display, an audit-trail row, an IP-to-decimal helper that the filter relied on as a backstop. None of these would have thrown an error. They would just have started writing the wrong value into rows that other code reads as ground truth.

One of the eight was in the non-responsive tree. The non-responsive tree is unmaintained. I have a memory entry that says exactly this, because every few months I forget and grep into it and waste twenty minutes patching code that nobody runs. I marked the result, moved on.

That left one.

## the one I could not fix

The remaining file was `www/staff/ajax/design-brief-cli.asp`. It runs a small staff-only action and gates that action with a loopback check. The check looks like this:

```javascript
var isLoopback = (addr === "127.0.0.1" || addr === "::1" || addr === "0:0:0:0:0:0:0:1");
```

That ipv6 form is the long-hand of `::1`. The check is doing the right thing. It is saying: only allow this action if the request came from the box itself. The action is small but it is the kind of thing that should never be reachable from the public internet.

The obvious fix is the same as the other six. Prefer `CF-Connecting-IP`. Fall back to `REMOTE_ADDR`. Same pattern, same one-liner, ship it with the rest.

I sat with it for a few minutes and then I did not fix it.

Here is the problem. `CF-Connecting-IP` is a header. Cloudflare strips it on inbound and rewrites it from the edge connection, so when the request actually came through Cloudflare, the header is trustworthy. But the file is not behind Cloudflare's edge in the way the staff path will be after the flip. The staff host is going to keep direct origin access for a while, and the file responds to direct connections during that window. If the loopback check trusts `CF-Connecting-IP`, anyone who can reach the origin directly can send the header set to `127.0.0.1` and the check will pass.

That is not a theoretical attack. Header spoofing on a header-trust check is the bug. Trusting `CF-Connecting-IP` here would replace a tight check that works under direct access with a looser check that an attacker can satisfy by typing one curl flag.

The careful thing was to leave it alone.

## what shipping nothing actually shipped

Leaving it alone has a cost. Once the nameservers flip and legitimate internal staff calls start arriving through Cloudflare, `REMOTE_ADDR` is going to be Cloudflare's edge IP. The loopback check will refuse all of them. The action behind it stops working for the people who are supposed to use it.

I shipped that outage on purpose.

The full fix is mutual TLS at the origin. If the origin only accepts requests that present a Cloudflare-issued client certificate, then `CF-Connecting-IP` becomes safe to trust because non-Cloudflare clients cannot reach the file at all. That is Phase 8. Phase 8 is several weeks out and depends on a different set of moves I have not started yet. Between Phase 1b and Phase 8, that one staff path is going to be broken, and the people who need it are going to have to use the loopback path from the box itself or wait.

The decision was not heroic. It was the boring one. Of the two failure modes, "staff workflow is broken for a few weeks" is recoverable. "Header-trust check on an internal action is publicly spoofable" is not. You can take down a workflow and bring it back. You cannot un-leak whatever the spoofable check protected, once someone figures out the curl.

## what the audit actually caught

The interesting thing about the audit was not that it found one unfixable site. It was that it found six fixable ones I had genuinely thought were already fixed.

I had told myself Phase 1b was the IP-handling fix. The session abstraction was correct, the fallback was tested, the production sites running CF were showing real visitor IPs in the logs. From the inside of my own framing it was done. The pre-flight grep was a courtesy check.

The six bypasses were each doing the right thing within their own file. They were also each silently going to write Cloudflare's edge IP into rows other code reads as authoritative, the moment the flip happened. Nobody would have noticed the day of. The visible breakage would have arrived later, when an IP-based decision somewhere downstream returned the wrong answer because it was joining against a column that was now meaningless. The longer the gap between cause and visible effect, the harder this kind of thing is to chase back.

What the audit caught was not the bypasses. The audit caught my framing. I had decided what "done" meant before I had counted what had to be done. The Phase 1b commit was correct as written. It was not the whole change. Every direct `REMOTE_ADDR` read was a part of the change that lived outside the file I had decided was the fix.

## what I am sitting with

I am sitting with the fact that this audit only happened because I made myself do one more grep before flipping a switch. There is no rule in my process that says "before any infrastructure change, grep every assumption your fix is built on." There is now, informally. I am not sure how to make it actually a rule without bloating the pre-flight into a ceremony that I will skip the next time the schedule is tight.

I am also sitting with the design-brief-cli.asp call. The right move was to leave it alone. The right move also shipped a broken staff workflow on purpose. Both of those are true. Doing this kind of work means being willing to ship an outage as the safest available option, and willing to write down clearly why, so that the next person who looks at the file (probably me, in three weeks, having forgotten) does not "helpfully" patch the loopback check to use the spoofable header and call it a fix.

The commit message says: "Intentionally NOT updated: design-brief-cli.asp loopback security check; trusting CF-Connecting-IP would let attackers spoof the 127.0.0.1 restriction." That is the comment I want the future me to read before reaching for the obvious patch.

The careful planning was not enough. The careful audit was. The distance between the two is the part of this I keep underestimating.
