---
title: "The customer's own scraper IS the bot swarm you've been fighting"
canonical: https://dxdev.com/blog/the-customers-own-scraper-is-the-bot-swarm/
datePublished: 2026-06-02
---
A customer opened a ticket asking us to allow-list their app's scraper so it could pull their own hosted site data. We were halfway through scoping the change when the obvious thing landed: that scraper was almost certainly the same AWS bot traffic that kept showing up pounding this exact account. The polite access request and the abuse incident were very likely the same traffic. The only thing that changed was the framing.

This is the trap I want to talk about, because it is easy to fall into and it sounds reasonable the whole way down. "It's the customer's own data" feels like it settles the security question. It does not. It is not a security argument at all. The load and the precedent are what matter, and neither one cares whose data is being pulled.

## How the ask arrived

The ticket started as the wrong question. Support's first answer, that we have no integrations, was answering a feature inquiry. The real ask only surfaced when the app's developer followed up: he runs AWS scrapers that crawl their hosted site to pull team, player, and stats data, and some of those requests were getting rate-limited. Translated, that is "your bot protection is rate-limiting my AWS scrapers, let them through."

That is a completely different request. The first is "can I connect to you in a supported way." The second is "stop applying the rule that is protecting every other site on the platform, specifically to me." When you hear the second one, the instinct to be helpful is exactly the instinct that gets you in trouble.

The account in question was a tournament org running on the platform. The scraper traffic originated from AWS, the same ASN family as the Bytespider and other AWS-swarm crawls our protection already catches. The platform runs system-wide bot protection precisely so that automated crawlers do not degrade the shared hosting for every tournament site at once. The customer's app was hitting that protection, getting rate-limited, and the developer wanted an exemption.

## Why "it's their data" doesn't hold

Walk the argument out and it falls apart in two places.

First, the load is real regardless of ownership. A scraper crawling for team, player, and stats data on a tight loop generates the same request volume whether the person behind it owns the team or not. The server does not get a discount for legitimate intent. The CPU, the connection pool, the database round-trips, those are all spent the same way. An account that keeps showing up getting pounded is getting pounded by request volume, and that volume does not become acceptable because we now know who is sending it. If anything, learning that the abuse traffic is the customer's own tool tells you the abuse will not stop on its own, because the tool is doing exactly what it was built to do.

Second, and this is the part that actually decides it: the precedent. The only mechanism available to "let AWS through" at the protection layer is blunt. You can carve out an IP, or you can carve out the AWS ASN. Carving out the ASN is the seductive option because it is one rule and it covers "all of AWS," which is where the customer's scraper lives. It is also a hole that every real scraper swarm on the internet drives straight through, because a huge fraction of hostile crawling already originates from AWS, GCP, and the other big clouds. Allow-list the ASN to satisfy one customer and you have effectively turned off cloud-origin bot protection for the whole platform. The next swarm that shows up from a different AWS account walks in through the door you propped open for this one.

So the exception you would build to be nice is the exact exception that undoes the protection you built to stay alive.

## What the right answer looks like

If you are going to make a scraper exception at all, and sometimes you legitimately should, the shape of it is fixed:

- Scope to exact egress IPs, never to an ASN. The app developer can tell you the specific addresses their scraper runs from. Those, and only those, go on the allow-list. A handful of /32s is a surgical hole. An ASN is the whole sky.
- Treat the IP list as something that drifts. Cloud egress IPs rotate. An allow-list of exact IPs is a maintenance commitment, not a fire-and-forget. That is a cost, and it is a fair reason to decline the whole thing if the customer cannot commit to stable egress.
- Recognize the request and the incident may be one event. Before you write the exception, get the scraper's actual egress IPs and check them against your abuse logs. If they match, you are not adding an integration, you are blessing the swarm. In our case we never had to collect the IPs to make the call, because the account was the one we kept seeing get pounded, and the source (AWS) lined up with the swarm family we already block. The pattern alone was enough to change the conversation.

In this case we declined, courteously. The platform has no public API and no programmatic export. The reply to the customer did not get clever about it. It explained that we block automated traffic system-wide to keep every hosted site fast and stable, and left it there. No code changed. The ticket closed as a customer reply, because the correct engineering action was to change nothing.

That is worth sitting with. The whole resolution was a decision not to build a feature, and the value was entirely in recognizing what the feature would have cost.

## The reframe that does the work


The reusable lesson is a reframe you can apply the next time a "polite access request" lands on your support queue.

Stop asking "is this person allowed to have this data?" That question almost always answers yes, because of course the customer can have their own roster. It is the wrong question and it leads you straight to the bad allow-list.

Start asking two other questions instead. What does granting this cost in load, sustained, not just in the demo? And what is the smallest possible blast radius of the mechanism I would use to grant it? If the only mechanism available is coarse (an ASN, a country, "all of cloud provider X"), then the answer is no until the customer can give you something narrow enough to scope to.

There is also a quieter discipline buried in here, which is that the source of a "feature request" and the source of an "incident" can be the same machine, and your ticket system will happily file them as two unrelated things. The recurring abuse on this account and the access request had different titles and different tones, and both pointed at AWS. The connection got made by recognizing the account, not by a forensic log dive, and that is the point: the cross-reference is cheap, often just "wait, isn't this the account that keeps getting hammered?", and it is the thing that flips the whole decision.

## Takeaway

"It's the customer's own data" is not a security argument. The load and the precedent are what decide it, and both are indifferent to ownership. Scope any scraper exception to exact egress IPs, never an ASN, and before you grant it, check your abuse logs, because the polite request to let a scraper in and the swarm you have been fighting are very often the same traffic wearing a nicer subject line.

## Related

- [The 3% bot attack that took the site down: why your IP blacklist can't see residential proxies](residential-proxy-swarm-blacklist-cant-see): why the same swarm family hides from standard IP reputation tools
- [The Prod Box Was DDoSing Itself: An iCal Calendar Feature Looping Through the Public Edge](prod-box-ddosing-itself-ical-through-public-edge): another case where "legitimate" traffic caused the same load damage as an attack
- [The UA blocklist is the right-sized fix for an identified crawler (and the wrong one for a swarm)](ua-blocklist-cheap-first-move-identified-crawler): choosing the right mitigation tier when the attacker is identified
- [The rate-limiter that banned our power users: 302 to 200 redirects look like an attack](rate-limiter-banned-power-users-302-200-redirects): a parallel case where protective rules fired on legitimate traffic
- [It Looked Exactly Like a Bot Swarm. It Was SQL Server Parameter Sniffing.](sql-server-parameter-sniffing-looked-like-a-bot-swarm): ruling out the swarm hypothesis before acting on it
