Running Scrapy behind a rotating pool
Why a single proxy breaks down first
A Scrapy spider running through one proxy looks fine in a test run and falls apart in production. The reason is concurrency. Scrapy is built to fire many requests in parallel, controlled by CONCURRENT_REQUESTS and CONCURRENT_REQUESTS_PER_DOMAIN. Every one of those parallel requests goes out through the same IP if you’ve hardcoded a single proxy in request.meta['proxy']. From the target site’s side, that looks like one address making an unusually high number of requests in a short window, which is exactly the pattern most rate-limiting and bot-detection systems are built to catch. A rotating pool exists to break that pattern by spreading requests across many exit IPs, so no single address carries the full weight of the crawl.
How Scrapy’s proxy middleware actually works
Scrapy ships with HttpProxyMiddleware, which sits in the downloader middleware chain at priority 750 by default. It’s simple: if a request has proxy set in request.meta, that value gets used as the upstream proxy for the request, read either from the meta dict directly or from http_proxy / https_proxy environment variables if meta isn’t set. That’s the entire built-in behavior. Scrapy doesn’t rotate anything on its own. Rotation is something you build on top, usually as a custom downloader middleware that runs before HttpProxyMiddleware in the chain and assigns a proxy to request.meta['proxy'] on every process_request call, pulling from a pool you maintain elsewhere.
The pool itself can be as simple as a list you round-robin through, or as involved as a small service that tracks proxy health and hands back a live one on request. Either way, the middleware’s job is narrow: pick a proxy, attach it to the request, and get out of the way. Retry and error handling live in separate middleware, which matters because proxy pools fail in specific, testable ways.
Sessions break rotation, and rotation breaks sessions
Rotating every single request sounds like the safest default, but it isn’t right for every crawl. If your spider logs in, holds a session cookie, and then paginates through results, switching IP on every request while keeping the same session cookie is its own kind of anomaly, since a site can see one session authenticated from ten different addresses inside a minute. For that kind of flow you want session stickiness: pin a request chain to one proxy for the life of the session, and only rotate when the session ends or the proxy starts failing.
A common pattern is to key your pool by a session ID and store the assigned proxy in spider.state or a custom dict, then have the rotation middleware look up the sticky proxy for that session ID instead of picking a fresh one on every call. Whether you rotate per-request or per-session should follow the shape of the site’s flow, not a fixed rule. Scraping a paginated listing with no login is a reasonable case for per-request rotation. Anything behind auth or a multi-step form usually isn’t.
Matching proxy type to what Scrapy is doing
Datacenter proxies are cheap and fast, and Scrapy’s async architecture can push a lot of throughput through them, but the IP ranges they come from are well documented and easy for a site to flag as hosting infrastructure rather than a residential ISP. They work best for targets that don’t do IP-reputation checks at all, or for internal and permissioned scraping where blocking isn’t a concern.
Residential proxies route traffic through real ISP-assigned addresses, which cost more per gigabyte and carry more latency variance, but they don’t carry the same “this is a data center” signal. That variance matters for Scrapy tuning: a pool of residential exits will have a wider spread of response times, so a fixed DOWNLOAD_DELAY tuned for datacenter speeds will either be too slow for the fast exits or too fast for the slow ones. AUTOTHROTTLE_ENABLED set to True handles this better than a static delay, since it adjusts request rate per-domain based on observed latency and error rate rather than a number you picked once.
Mobile proxies route through carrier networks and share IPs across many real mobile users behind carrier-grade NAT, which is a different kind of cover than a residential IP does. They’re the most expensive tier and the slowest to provision at volume, so they make more sense for a narrow set of hard targets than as a default for a whole crawl.
None of these tiers are undetectable, and none of them guarantee a clean run. They change the odds and the cost profile, not the outcome.
Handling proxy failure without hiding it
A live pool has dead and slow proxies in it constantly, and Scrapy needs to know the difference between “the target site blocked this request” and “this proxy is unreachable.” Those get handled in different places. Connection failures and timeouts raise exceptions that process_exception in your middleware can catch, letting you drop that proxy from rotation and retry the request on a different one. HTTP-level responses that indicate blocking, like 403s, 429s, or a 200 that actually contains a challenge page, don’t raise exceptions at all since as far as the transport is concerned the request succeeded. Those need to be caught in process_response, either in your own middleware or by extending RetryMiddleware with RETRY_HTTP_CODES set to include them.
It’s worth logging both categories separately with the proxy IP attached. A pool where the same handful of proxies keep timing out is a supplier problem. A pool where every proxy eventually gets a 403 on the same domain is a targeting problem, and swapping proxies faster won’t fix it. Conflating the two makes debugging a stalled crawl much harder than it needs to be.
What detection is actually looking at
Sites that want to identify automated traffic aren’t only looking at IP reputation. Request headers matter: a default Scrapy user agent, or one that doesn’t match the rest of the header set a real browser would send, is one of the easier tells. Timing matters too, since real users don’t request pages at perfectly even intervals, which is part of why AUTOTHROTTLE_ENABLED and some jitter on delay produce a less mechanical pattern than a fixed delay does. At a more advanced level, some detection systems fingerprint the TLS handshake itself, since Scrapy’s default HTTP client has a distinct TLS signature from a real browser’s. None of this is a checklist for getting through defenses undetected. It’s the actual set of signals a site’s operators are reasonably justified in watching for, because those same signals are what separates a legitimate crawler from a credential-stuffing bot or a scraper harvesting data the site never intended to expose in bulk.
Keeping a crawl clean at scale
The practices that keep a scrape stable over months are unglamorous. Read robots.txt and respect ROBOTSTXT_OBEY where the target’s terms call for it. Set a real, identifying user agent rather than spoofing a browser string if the site allows crawler traffic, since a false UA turns a routine block into a trust problem if anyone looks closer. Keep concurrency and request rate proportional to what the target can absorb, not the maximum your pool can technically push. Scope the crawl to public, non-authenticated data unless you have explicit permission for anything behind a login. And build monitoring into the pool from day one: success rate per proxy, error codes per domain, and rotation frequency, so a degrading pool or a tightening target shows up as a metric before it shows up as a crawl that silently stopped collecting data three days ago.
None of this makes a scraper unblockable. Sites change their detection, pools include a bad proxy from time to time, and a crawl that ran clean last month can trip something new this month. The goal of a well-built middleware setup is to fail visibly and recover automatically when that happens, not to promise it won’t.
If you’re comparing proxy pools for a Scrapy project, or want an honest look at how residential, mobile, and datacenter options actually behave under real crawl traffic, head to the Proxy Scraping homepage for the rest of our testing and guides.
Get new guides and videos first — join the Telegram channel.