< Back

Managing AI Crawler Traffic on Your Website: How Reverse Proxies and IP Reputation Filtering Control GPTBot and Other LLM Scrapers

Tech

A content team I spoke with last year opened their CDN dashboard after a routine release and found that non-human traffic had quietly become the majority of their requests. Not search engines. Not uptime monitors. A mix of declared AI training crawlers, retrieval agents fetching pages on behalf of chat users, and a long tail of unidentified clients hammering their faceted product filters at three requests per second per IP.

Nothing was down. Nothing was obviously broken. But origin egress had nearly doubled, their cache hit ratio had fallen off a cliff, and the analytics team was reporting "traffic growth" that converted at zero percent.

This is the position most publishers, SaaS marketing sites, and e-commerce catalogs are now in. AI crawlers are a genuinely new traffic class with different economics from search bots: they fetch deeply, they fetch often, they rarely send referral traffic back in proportion to what they take, and a meaningful share of them do not identify themselves honestly. Deciding what to do about that is no longer an abstract policy debate. It is an infrastructure decision, and the place it gets made is your reverse proxy.

Who Is Actually Crawling You

Before you write a single rule, you need to separate the traffic into categories that behave differently and deserve different treatment. Lumping everything under "AI bots" produces bad policy.

Declared training and indexing crawlers

These are the well-behaved ones. They publish a user agent string, they publish IP ranges, and they read robots.txt. The familiar names include GPTBot, ClaudeBot, CCBot (Common Crawl, which feeds many downstream models), Google-Extended (a robots.txt token rather than a distinct crawler), Applebot-Extended, Meta-ExternalAgent, Amazonbot, and Bytespider.

They are also the easiest to control, which is why teams overweight them. Blocking every declared crawler feels decisive and changes very little about your actual load, because the declared bots were rarely the ones causing the damage.

Live retrieval agents

This category matters more than most site owners realise. When a user asks a chat assistant a question and the assistant fetches a live page to answer it, that request often carries a distinct identity: ChatGPT-User, Perplexity-User, OAI-SearchBot and similar. The fetch is user-initiated, single-page, and low-volume.

Those requests are the closest thing to a click that the AI answer layer currently produces. If your blanket rule catches them alongside the bulk training crawlers, you have removed yourself from the answer surface that is slowly replacing part of your organic search funnel. Treat retrieval agents as a distinct policy bucket from training crawlers, because they represent distribution, not extraction.

Undeclared and disguised scrapers

The third category is the one that actually consumes your budget. These clients present a common Chrome user agent, run headless browsers, rotate IP addresses aggressively, ignore robots.txt entirely because they never read it, and often originate from commodity cloud ranges or from residential IP space rented by the gigabyte.

Some of this is AI data collection by intermediaries who resell datasets. Some is competitive price monitoring. Some is straightforward content theft. From your edge, they are indistinguishable by name, which is exactly the point: the only signals you have are behaviour, network provenance, and fingerprint consistency.

robots.txt Is a Policy Statement, Not a Control

robots.txt remains worth maintaining. It is the legally and reputationally useful record of your stated preference, it is honoured by the major declared crawlers, and increasingly it is the artefact that licensing conversations start from.

What it is not is enforcement. It is a text file requesting cooperation. Every actor that already decided to ignore your terms will continue to ignore them, and the gap between your robots.txt and your access logs is precisely the traffic that needs an infrastructure answer.

There is also a subtler failure mode: teams write a sprawling robots.txt with dozens of disallow directives, never validate it against real logs, and assume compliance. Compliance is measurable. Grep your logs for the declared crawler that you disallowed last quarter and confirm it actually stopped. Occasionally it did not, because the directive was scoped to a path that redirects, or because a second user agent from the same operator was never covered.

The Reverse Proxy as the Control Plane

The right place to enforce crawler policy is the layer that sees every request before your application does: Nginx, Envoy, Caddy, HAProxy, or your CDN's edge worker runtime. Doing it in application middleware is tempting and almost always wrong, because by then you have already paid for TLS termination, routing, database connections, and template rendering.

A reverse proxy gives you four capabilities that application code cannot match cheaply.

Single point of policy. One rule set applied consistently across every hostname, path, and backend. No drift between your marketing site, your docs, and your API.

Cheap rejection. Returning a 403 or 429 at the edge costs you a handful of microseconds and a few hundred bytes. Rendering a page for a crawler that will never convert costs you a database round trip.

Rate shaping instead of binary decisions. Token bucket limiters keyed on verified identity, ASN, or IP let you say "you may crawl, at two requests per second, with a burst of ten" rather than "yes" or "no".

Observability. Structured logs that record the classification decision alongside the request give you the dataset you need to tune. Without it you are guessing.

Verify declared crawlers properly

The single most common implementation mistake is trusting the user agent string. Anyone can send User-Agent: GPTBot. In practice, plenty of low-effort scrapers do exactly that, because they have learned that some sites whitelist known crawlers to preserve SEO.

There are two defensible verification methods. The first is forward-confirmed reverse DNS: take the client IP, resolve the PTR record, confirm it belongs to the operator's domain, then resolve that hostname forward and confirm it returns the original IP. The second, which most AI crawler operators now support, is matching the client IP against the published IP range list that the operator hosts as a JSON file.

The practical approach is to fetch those published ranges on a schedule, compile them into an IP set your proxy can match in constant time, and treat an unverified crawler user agent as a stronger negative signal than an anonymous one. A client claiming to be GPTBot from an IP outside the published range is not a mislabelled browser. It is lying to you.

A simplified Nginx sketch of the classification idea:

geo $ai_verified {
default 0;
include /etc/nginx/ai-ranges/openai.conf; # generated hourly
include /etc/nginx/ai-ranges/anthropic.conf;
}

map $http_user_agent $ai_claimed {
default 0;
"~*GPTBot|ClaudeBot|CCBot|Bytespider" 1;
}

# claimed but not verified: reject early
if ($ai_claimed$ai_verified = "10") { return 403; }

limit_req_zone $binary_remote_addr zone=ai_slow:10m rate=2r/s;

Keep the real logic in a worker or a WAF rule set where you can express it cleanly, but the shape holds: claim, verification, then rate class.

Shape the traffic instead of banning it

Blanket blocking is a blunt instrument that invites an arms race. The operator who respects your robots.txt today and gets a 403 tomorrow has little incentive to keep identifying itself. Meanwhile the scrapers that never identified themselves are unaffected.

Graduated response works better in production. Verified training crawlers get a low, steady rate limit and access only to the sections you are happy to have ingested. Retrieval agents get a normal rate limit because they arrive one page at a time on behalf of a human. Unverified clients with crawler-like behaviour get a challenge. Confirmed abusive patterns get a hard block at the edge.

Fix the crawl surface, not just the crawlers

A large share of AI crawler cost is self-inflicted. Faceted navigation that generates millions of URL permutations, calendar widgets with infinite next-month links, session identifiers in query strings, and search result pages that are indexable all create effectively unbounded crawl space. Any crawler that finds it will consume it.

Canonicalise aggressively, return 404 or 410 for permutations you do not want fetched, strip tracking parameters at the edge before cache lookup, and make sure your cache key does not fragment on parameters that do not change the response. Teams routinely cut AI crawler origin load by half through cache key hygiene alone, before touching a single bot rule.

IP Reputation Filtering: What It Can and Cannot Do

IP reputation is the second control layer, and it is genuinely useful as long as you understand what it actually measures.

A reputation score is an aggregate of network provenance and observed history. The strongest inputs are ASN classification (is this IP in hosting space, consumer broadband space, or mobile carrier space), historical abuse reports tied to the address or its neighbours, how recently the range was allocated or re-allocated, whether the IP appears in known open proxy and VPN inventories, and what other traffic from the same /24 has been doing across the network that computes the score.

What it does well: cloud and hosting ASNs are a strong signal. A request claiming to be a consumer Chrome browser on Windows, arriving from a virtual machine range in a datacenter, with no prior session history on your site, is almost never a real customer. Scoring that heavily is cheap and low-risk.

What it does badly: residential and mobile space. Carrier-grade NAT means a single mobile IP can front thousands of real subscribers. Corporate VPN egress concentrates entire companies behind one address. University networks, ISP proxies, and shared office ranges all produce reputation noise. If you block on low residential reputation alone, you will block paying customers, and you will not find out from your logs because those users simply leave.

Three rules keep reputation filtering honest.

Score, do not gate. Reputation should be one weighted input into a composite decision that also includes request rate, path entropy, header ordering, TLS fingerprint consistency, and whether the client executes JavaScript. No single signal should be able to produce a block on its own.

Escalate proportionally. Low score plus benign behaviour earns a lightweight challenge. Low score plus crawl-shaped behaviour earns a rate limit. Only sustained, clearly automated abuse earns a hard block.

Measure your false positives. Instrument challenge pass rates by ASN and by country. If a specific mobile carrier in a market you sell into is failing challenges at an unusual rate, that is a revenue bug, not a security win.

Common Mistakes Teams Make

Blocking by ASN wholesale. Banning an entire cloud provider's ASN also bans your uptime monitors, your partner integrations, your own CI pipeline, accessibility tooling, and any legitimate service that calls you server to server. Scope cloud ASN rules to browser-claiming clients only.

Killing citation traffic by accident. Grouping live retrieval agents with training crawlers is the most expensive avoidable error in this space right now.

Serving different content to different bots without thinking it through. If your edge rules end up returning a substantially different page to a crawler than to a human, you are in cloaking territory, with search consequences that dwarf whatever you saved on bandwidth.

Leaving the side doors open. Teams lock down HTML and forget the JSON API the frontend calls, the RSS feed, the sitemap that enumerates every URL, and the staging subdomain with no rules at all.

Treating it as a one-time project. User agents change, IP ranges rotate, new agents appear monthly. Crawler policy is a maintained system with an owner, not a ticket you close.

Where Proxies Fit In

There is an obvious irony in a proxy discussion appearing in an article about blocking automated traffic, but the two sides of this problem are the same discipline viewed from opposite ends, and you cannot do the defensive side well without the offensive toolkit.

Once you deploy classification rules at your edge, you have created a set of conditional behaviours that your own monitoring cannot see. Your uptime checks come from a fixed set of datacenter IPs that you almost certainly whitelisted. Your team browses from the office or a corporate VPN. Nobody on the inside is experiencing the rules the way a customer on a mobile carrier in Brazil or a residential line in Poland experiences them.

The only reliable way to audit that is to request your own site from the same kinds of networks your users actually sit on. Testing from diverse residential and mobile proxy pools across the markets you sell into tells you whether your challenge rate is uniform or whether one carrier, one ASN, or one country is quietly being penalised. Run the same set of URLs through a residential exit, an ISP exit, a mobile exit, and a datacenter exit, and compare status codes, challenge frequency, and time to first byte. The differences are your policy, rendered visible.

This is where pool composition matters more than raw volume. Auditing a reputation-based rule set requires exit nodes that genuinely represent consumer networks rather than repackaged hosting space, and it requires enough geographic spread to cover the regions where your revenue lives. EnigmaProxy positions itself in the professional tier on exactly those axes: multiple pool types (residential, ISP, datacenter, and mobile) under one account, broad geo-coverage for regional verification, ethical sourcing of residential nodes, and session control that lets you hold an IP long enough to walk through a multi-step flow rather than rotating mid-journey and invalidating your own test. When a single result looks anomalous, it helps to be able to check an individual exit IP before you conclude that your edge rules are misconfigured.

The same infrastructure serves the other half of the equation. If your organisation also collects public web data, the standards you want applied to you are the standards you should apply outward: honour robots.txt, identify your agent where appropriate, hold concurrency to a level that does not degrade the target, and run on ethically sourced pools rather than networks of uncertain provenance. Publishers are getting better at telling the difference, and the collectors who behave well are the ones who keep their access.

Where This Is Heading

Machine-readable licensing becomes the default. robots.txt was built to express yes or no. What the market now needs is price, permitted use, and attribution terms, expressed in a format an agent can parse. Emerging content licensing standards and the revival of HTTP 402 as a real status code both point the same direction: crawl requests that carry payment or licence context, negotiated at the edge rather than in a legal department.

Identity replaces IP as the primary trust signal. Cryptographically signed bot requests, where an agent proves who it is with a key rather than a string, are already being piloted. That will not eliminate IP reputation, but it will demote it from primary evidence to corroborating evidence, and it will make the undeclared long tail stand out far more sharply than it does today.

Agentic commerce changes the cost calculation. When an AI agent is not just reading your catalog but completing a purchase on a user's behalf, blocking it stops being a bandwidth saving and starts being a lost sale. Expect the policy question to shift from "how do we keep bots out" to "which bots are customers and how do we authenticate them".

Analytics gets rebuilt around the answer layer. Attribution models that assume a click cannot measure a channel where the assistant summarises your page and the user never arrives. Server-side log analysis of retrieval agent fetches, correlated against branded search lift, is becoming a real reporting discipline rather than a curiosity.

Closing Thoughts

AI crawler traffic is not a problem to be solved once. It is a new class of network participant that needs a policy, an enforcement layer, and a measurement loop, in that order.

Write the policy in robots.txt because it is the record that matters. Enforce it at the reverse proxy, where rejection is cheap and rules are consistent. Verify declared crawlers against published ranges instead of trusting strings. Use IP reputation as a weighted signal within a composite decision, never as a standalone gate, and watch your false positive rate as closely as your block rate. Then fix the crawl surface, because the cheapest request is the one your architecture never invited.

And audit the whole thing from outside your own network, repeatedly, from the kinds of connections your customers actually use. For teams that need dependable geo-diverse exit nodes to run that verification with business-grade reliability, EnigmaProxy is a solid option to evaluate alongside the rest of your edge tooling.