< Back

Async Python Web Scraping with Rotating Proxies: Combining asyncio and aiohttp for High-Throughput Pipelines in 2026

Tutorials

A scraper that pulls 40,000 product pages with Python's requests library and a loop will spend roughly 95 percent of its runtime doing nothing at all. The CPU idles while the socket waits for bytes to come back over the wire. Add a proxy hop and a TLS handshake to each request and the dead time grows further. That is the whole problem async solves: not faster requests, but far more of them in flight at once.

The catch is that concurrency multiplies every weakness in your proxy layer. A rotation strategy that looks fine at five requests per second starts producing 403s, connection resets, and half-written rows the moment you push it to five hundred. Teams usually discover this the hard way, somewhere around the third week of production, when the dataset quietly starts missing a region.

This guide covers how to build an asyncio and aiohttp scraping pipeline that actually holds up under load: how aiohttp handles proxies at the protocol level, how to structure rotation so it survives concurrency, how to apply backpressure, and which mistakes silently corrupt data rather than throwing errors.

Why asyncio Changes the Throughput Math

Synchronous HTTP is blocking I/O. One thread, one request, one wait. Threading improves on that but each thread costs memory and context switching, and the GIL means you are juggling OS threads to work around an interpreter limitation rather than solving the real issue.

asyncio replaces threads with an event loop and coroutines. A single thread can hold thousands of open sockets, switching between them whenever one is waiting on the network. For scraping workloads, which are almost entirely network-bound, this is close to ideal. The practical ceiling moves from "how many threads can I afford" to "how many concurrent connections will my proxy pool and the target site tolerate".

That second constraint is the one most engineers underestimate. The event loop will happily open 2,000 sockets. Your exit nodes, your target's rate limiter, and your local file descriptor limit will not all agree with that decision.

Where aiohttp Sits

aiohttp is the mature async HTTP client for Python, and in 2026 it remains the default for high-volume pipelines. httpx offers a nicer API and HTTP/2 support, which matters for some targets, but aiohttp's connector model gives you finer control over connection pooling and per-host limits, and that control is exactly what proxy-heavy workloads need.

The key object is ClientSession. It owns a connection pool, cookie storage, and default headers. Creating a session per request is one of the most common performance bugs in async scrapers: you throw away every keep-alive connection and pay a fresh TCP plus TLS handshake every time, which through a proxy can mean 300ms of pure overhead per call.

How aiohttp Actually Handles Proxies

aiohttp supports HTTP and HTTPS proxies natively through the proxy parameter on each request, plus proxy_auth for credentials:

import aiohttp

async def fetch(session, url, proxy_url):
async with session.get(
url,
proxy=proxy_url,
timeout=aiohttp.ClientTimeout(total=25),
) as resp:
return resp.status, await resp.text()

If your proxy uses user and password authentication, you can embed credentials directly in the URL (http://user:[email protected]:8000) or pass a BasicAuth object. Embedding is fine and widely used, but be careful: credentials in URLs tend to end up in logs. If you log failed requests with their proxy string, you are writing passwords to disk.

Two protocol details matter here.

HTTPS targets go through CONNECT. When you request an https:// URL via an HTTP proxy, aiohttp issues a CONNECT tunnel and performs the TLS handshake end to end with the target. The proxy sees the hostname via SNI but not the payload. This is normal and correct, but it means each new tunnel is expensive. Reusing connections through the same exit IP is significantly cheaper than rotating on every single request.

SOCKS5 needs an extra package. aiohttp does not speak SOCKS natively. Install aiohttp-socks and pass a ProxyConnector to the session instead of using the per-request proxy argument. That is an important structural difference: with SOCKS the proxy is bound to the connector, so per-request rotation requires either a gateway that rotates server-side or multiple connectors. For most scraping work, an HTTP proxy endpoint with rotation handled upstream is the simpler path.

Designing the Pipeline

A production async scraper has four moving parts: a work queue, a bounded set of workers, a fetch layer with retries, and a sink that writes results. Keeping them separate is what lets you tune throughput without rewriting logic.

Bound Concurrency in Two Places

Use a semaphore for logical concurrency and a connector limit for socket-level concurrency. They are not the same thing and you need both.

import asyncio
import aiohttp

CONCURRENCY = 120

connector = aiohttp.TCPConnector(
limit=CONCURRENCY,
limit_per_host=20,
ttl_dns_cache=300,
enable_cleanup_closed=True,
)

sem = asyncio.Semaphore(CONCURRENCY)

limit_per_host is the one people forget. Without it, 120 concurrent workers can all hammer the same origin through the same proxy gateway and trip a rate limiter that would otherwise never have fired. Setting it to a sane fraction of total concurrency spreads load naturally.

ttl_dns_cache matters more than it looks. Under high concurrency, uncached DNS lookups become a bottleneck and, worse, a fingerprint: thousands of identical lookups from one resolver is a pattern.

Producer, Workers, Sink

async def worker(name, queue, session, rotator, results):
while True:
url = await queue.get()
try:
data = await fetch_with_retry(session, url, rotator)
if data:
await results.put((url, data))
finally:
queue.task_done()

async def run(urls):
queue = asyncio.Queue(maxsize=5000)
results = asyncio.Queue(maxsize=5000)
rotator = ProxyRotator(load_proxies())

async with aiohttp.ClientSession(connector=connector) as session:
workers = [
asyncio.create_task(worker(i, queue, session, rotator, results))
for i in range(CONCURRENCY)
]
writer = asyncio.create_task(write_results(results))

for url in urls:
await queue.put(url)

await queue.join()
for w in workers:
w.cancel()
await results.put(None)
await writer

The bounded maxsize on both queues is the backpressure mechanism. If your writer falls behind (a slow database, a saturated disk), the results queue fills, workers block on put, and the fetch rate drops automatically. Unbounded queues let a scraper consume all available memory while appearing perfectly healthy right up until the OOM kill.

Rotation Strategies That Survive Concurrency

This is where async scraping diverges sharply from the synchronous version. In a sequential script, "rotate the proxy each request" is unambiguous. With 120 coroutines in flight, you need to decide what rotation means when requests overlap.

Gateway Rotation vs Client-Side Rotation

With a rotating gateway endpoint, you point every request at one host and port, and the provider assigns a different exit IP per connection. Your code stays trivially simple and your concurrency is limited only by what the gateway allows. This is the right default for broad crawls where any IP in the target country will do.

Client-side rotation means you hold a list of endpoints or sticky session identifiers and pick one per request. You get explicit control: you can pin a session to a cart, a login, or a paginated sequence, and you can retire an IP after a specific failure. The cost is that you now own the health tracking.

A Rotator With Health Tracking

A list and random.choice() is not a rotation strategy. It will keep handing out IPs that are already blocked on your target, and under concurrency those failures compound fast.

import time, random
from collections import defaultdict

class ProxyRotator:
def __init__(self, proxies, cooldown=90, fail_threshold=3):
self.proxies = list(proxies)
self.cooldown = cooldown
self.fail_threshold = fail_threshold
self.failures = defaultdict(int)
self.benched_until = {}

def get(self):
now = time.monotonic()
live = [
p for p in self.proxies
if self.benched_until.get(p, 0) <= now
]
if not live:
live = self.proxies
return random.choice(live)

def report(self, proxy, ok):
if ok:
self.failures[proxy] = 0
return
self.failures[proxy] += 1
if self.failures[proxy] >= self.fail_threshold:
self.benched_until[proxy] = time.monotonic() + self.cooldown
self.failures[proxy] = 0

Benching rather than permanently removing matters because most block signals are temporary. An IP that returns 429 for ninety seconds is usually fine afterwards. Discarding it permanently shrinks your usable pool over a long run until throughput collapses for reasons nobody can explain.

Sticky Sessions for Stateful Flows

When a target requires a login, a currency selection, or a multi-step pagination token, the exit IP must stay constant for the duration. Most providers expose this through a session identifier in the username field. Structure your work so that one coroutine owns one session for the whole flow, and never let the generic rotator touch it. Mixing sticky and rotating traffic through one code path is how you get cart contents that belong to another worker.

Retries, Backoff, and an Error Taxonomy

Treating all failures identically is the single biggest cause of wasted bandwidth in async scrapers. Different errors demand different responses.

Connection errors and timeouts usually mean a bad exit node, not a block. Retry immediately on a different proxy, up to a small limit. These should not count heavily against your pool health unless one IP produces them repeatedly.

429 and 503 mean you are going too fast. Retrying instantly on a new IP works briefly and then stops working, because the target is often rate limiting by fingerprint or subnet as well as by IP. Back off with jitter and reduce global concurrency temporarily.

403 and challenge pages mean the request was identified. Rotating the IP alone rarely fixes this. Check that your headers, TLS profile, and IP type are consistent before burning more of the pool.

200 with the wrong content is the dangerous one. Soft blocks return a valid status code and a page that contains a consent wall, an empty result set, or a generic "we noticed unusual traffic" notice. Validate content, not status codes.

import asyncio, random, aiohttp

RETRYABLE = {408, 425, 429, 500, 502, 503, 504}

async def fetch_with_retry(session, url, rotator, attempts=4):
for attempt in range(attempts):
proxy = rotator.get()
try:
async with session.get(
url, proxy=proxy,
timeout=aiohttp.ClientTimeout(total=25, connect=8),
) as resp:
body = await resp.text()
if resp.status == 200 and looks_valid(body):
rotator.report(proxy, True)
return body
rotator.report(proxy, False)
if resp.status not in RETRYABLE and resp.status != 403:
return None
except (aiohttp.ClientError, asyncio.TimeoutError):
rotator.report(proxy, False)
await asyncio.sleep((2 ** attempt) + random.random())
return None

The jittered exponential backoff is not decoration. Without jitter, a burst of failures causes all retries to fire at the same instant, producing a thundering herd that the target reads as an attack pattern.

Common Mistakes in Async Proxy Pipelines

Running blocking code inside coroutines. One synchronous database write, one requests.get, or one unoptimised BeautifulSoup parse of a 3MB document will freeze the entire event loop. Heavy parsing belongs in run_in_executor or a separate process pool.

Unlimited concurrency against one host. asyncio.gather over 10,000 URLs opens 10,000 tasks. Even with a connector limit, the task overhead and memory footprint are real, and the moment your pool is saturated every request starts timing out at once.

Ignoring bandwidth, not just request count. Async pipelines pull data fast. On a metered residential plan, an unthrottled image-heavy crawl can consume a month of budget in a day. Set max_field_size, skip binary responses you do not need, and request compressed encodings.

No per-proxy metrics. If you cannot answer "which exit nodes produced our failures last night", you are flying blind. Log proxy identifier, status, latency, and response size per request. Aggregate hourly.

Assuming a successful write means a successful scrape. Count validated records, not completed tasks. A pipeline that writes 40,000 rows of which 6,000 are soft-blocked placeholders is worse than one that writes 34,000 and tells you about the gap.

Where Proxies Fit In: Matching Pool Type to Async Throughput

Async code can only go as fast as the network path underneath it allows. Once you remove the blocking bottleneck from your own process, the proxy layer becomes the limiting factor, and the characteristics that matter change.

Concurrency tolerance comes first. Some pools behave well with a handful of simultaneous connections and degrade sharply beyond that. A high-throughput pipeline needs an endpoint that accepts hundreds of parallel sessions without queueing them, which is largely a function of pool size and gateway architecture rather than raw speed.

Pool diversity comes second. Datacenter IPs are the cheapest way to move volume and work well against tolerant targets and public APIs. ISP proxies give you static, residential-registered addresses with datacenter latency, which suits long sticky sessions. Rotating residential proxy pools are what you reach for when the target scores IP type directly, and mobile pools sit behind carrier-grade NAT, making them the most resilient option for the hardest endpoints. Serious pipelines mix all four and route by target class rather than paying residential rates for everything.

Geo-coverage is the third axis. Localised pricing, regional catalogues, and country-gated content all require exit nodes in the right market, and "available in 150 countries" means little if the city-level availability in the markets you actually need is thin. EnigmaProxy operates multiple pool types with ethically sourced residential IPs and broad geo-coverage, which is the kind of profile that matters when a single crawl needs German residential IPs for one segment and datacenter capacity for another in the same run.

Predictable cost is the part that only becomes obvious at scale. Async scrapers consume bandwidth in bursts, so a pricing model you can forecast against is worth more than a low headline rate you cannot plan around. Reviewing EnigmaProxy plans against your measured gigabytes per million requests is a more useful exercise than comparing per-GB numbers in isolation.

Before a large run, it is also worth confirming that the endpoints you intend to use resolve, authenticate, and geolocate where you expect. A quick pass through a proxy tester catches misconfigured credentials and unexpected exit locations in minutes, which is considerably cheaper than discovering them three hours into a crawl.

Strategic Shifts to Plan For

Detection is moving to behaviour and timing. Anti-bot systems increasingly profile request cadence, concurrency patterns, and the statistical regularity of inter-request gaps. Perfectly uniform async traffic is itself a signal. Expect to add deliberate variance and per-session pacing rather than running every worker at maximum speed.

HTTP/2 and HTTP/3 adoption changes the fingerprint surface. More targets now expect modern protocol behaviour, and clients that negotiate HTTP/1.1 against an HTTP/2-first origin stand out. This is pushing some teams toward clients with native HTTP/2 support, with proxy compatibility becoming a selection criterion rather than an afterthought.

Hybrid pipelines are becoming standard. Running everything through a headless browser is slow and expensive; running everything through raw HTTP fails on JavaScript-heavy pages. The emerging pattern is an async HTTP tier handling 90 percent of volume, with a small browser tier reserved for pages that genuinely need rendering, both drawing from the same proxy pool with consistent geo-targeting.

Observability is becoming a first-class requirement. Teams that treat proxy metrics the way they treat application metrics (success rate by pool, latency percentiles by region, bandwidth per thousand records) catch degradation days earlier. In 2026 this is increasingly the difference between a pipeline that is trusted by the business and one that is quietly distrusted.

Conclusion

asyncio and aiohttp give Python scrapers the throughput to collect at serious scale, but the gains are only real if the surrounding architecture respects what concurrency does to everything else. Reuse sessions. Bound concurrency at both the semaphore and the connector. Classify errors instead of blanket-retrying them. Track proxy health rather than picking at random. Validate content rather than trusting status codes. Apply backpressure through bounded queues so a slow sink throttles the crawl instead of exhausting memory.

The proxy layer deserves the same engineering attention as the code. Pool type should match target difficulty, geo-coverage should match the markets in your dataset, and sourcing should be something you can explain to a compliance reviewer. Providers such as EnigmaProxy, with multiple pool types and a focus on ethical sourcing and business-grade reliability, fit that brief for teams running high-throughput pipelines where data quality is the actual deliverable.