← all guides

Choosing a concurrency level your proxies can sustain

Most scraping jobs don’t get blocked because someone found a clever detection trick. They get blocked because the concurrency setting was picked to match the scraper’s CPU, not the proxy pool underneath it. If you run 500 threads against a target and only have 80 proxies rotating behind them, every proxy is carrying more than six simultaneous connections, and that pattern is visible to the target long before anyone inspects headers or TLS fingerprints.

This piece is about the sizing problem: how to pick a concurrency level that your proxy pool can actually sustain, and why that number depends on the pool, the proxy type, and the target’s own traffic shape, not on how many threads your machine can spin up.

What concurrency means in this context

Concurrency is the number of requests your scraper has in flight at the same time, whether that’s threads, async tasks, or browser contexts. It’s easy to confuse with requests per second, but they’re not the same thing. A job doing 20 concurrent requests against a slow endpoint that takes 4 seconds to respond is only pushing about 5 requests per second. The same 20 concurrent slots against a fast endpoint that responds in 200ms pushes 100 requests per second. Concurrency is the dial you set; request rate is what comes out the other end, and it depends on target latency as much as your settings.

This matters because proxy pools are consumed by concurrency, not by request rate directly. Each concurrent request needs an IP to go out on. If your pool has fewer usable IPs than your concurrency setting, IPs get reused simultaneously, and that reuse is the thing that creates a traffic pattern a target can notice.

Why this is a proxy problem, not just a scraper setting

Every proxy is standing in as an identity in front of the target site. A residential or mobile IP represents, from the target’s point of view, something like a single household connection or a single carrier subscriber. A datacenter IP represents a block of infrastructure that the target’s systems can often recognize by ASN.

When your concurrency exceeds what a given proxy or pool can carry gracefully, you get parallel requests landing from the same identity at the same moment, sometimes hitting different paths on the same site within milliseconds of each other. A real single visitor doesn’t load ten unrelated pages of a site in parallel. Rate limiters and WAFs are built to notice exactly this kind of clustering, because it’s one of the more reliable signals that a request isn’t coming from an ordinary browser session.

Sizing the pool against the job

The working relationship is: safe concurrency is a function of how many usable proxies you have and how much load each one can take before its behavior looks abnormal for that proxy type. You can write it loosely as:

concurrency ceiling ≈ (number of active proxies) × (safe simultaneous requests per proxy)

The second term is the part providers rarely publish and targets never document, so you have to infer it by watching your own error rates. A reasonable way to find it without guessing is to start low, for example one request in flight per proxy at a time, run the job, and watch for the failure signals described further down. If none show up, increase in small steps and keep watching. This is closer to how autothrottle style tooling in Scrapy already works: it starts conservative and adjusts based on response latency and error rate, rather than assuming a fixed number.

Datacenter, residential, and mobile carry load differently

Datacenter proxies are usually the cheapest and the most consistent in latency, but they sit on ranges that are easy for a target to attribute to hosting infrastructure at the ASN level. The risk with datacenter pools isn’t usually one IP being hit too hard, it’s the whole subnet getting a reputation hit if enough of the pool’s IPs draw attention around the same time. Concurrency planning here has to think about the pool’s aggregate footprint on the target, not just per-IP load.

Residential proxies route through real ISP-assigned connections. A single residential IP handling two or three simultaneous connections isn’t inherently strange, since that’s what a household with a couple of open browser tabs looks like. What is strange is that same IP hitting many different, unrelated resources on a target in a short window with no idle time between requests, which is a shape normal browsing rarely produces. Residential pools tend to tolerate a bit more per-IP concurrency than datacenter, but they still have limits, and providers vary a lot in how they manage IP churn and reuse.

Mobile proxies route through carrier networks, and carrier-grade NAT already means many real subscribers can share the same public IP at once. That gives mobile pools some natural cover for concurrent traffic, but it cuts both ways: carriers also reassign these IPs, and a mobile IP that inherits a bad reputation from a previous user is a real operational cost, not a hypothetical one. Rotation cadence matters more here than raw concurrency per IP.

The rotation interval trap

A common assumption is that rotating to a new IP on every single request is the safest option. It isn’t automatically true, and it can work against you. Many detection systems are built around session coherence: they expect a sequence of requests from one identity to look like a continuous visit, with cookies carried forward, a consistent user agent, and a plausible pace between pages. If your concurrency is high and your rotation interval is aggressive, you end up sending a wall of single-request, no-history connections that all start cold at once. That’s its own distinct pattern, and it’s not obviously better than a slower, session-consistent one.

Concurrency and rotation interval need to be tuned together. A sticky session held for a reasonable span with humanlike pacing between requests is often a better match for how session-based detection is built than pure per-request rotation at high concurrency.

Reading the signals that you’re over the ceiling

You don’t need to see a target’s internal rate-limiting logic to know when you’ve exceeded it. The job tells you:

  • 403 and 429 response rates climbing as the run continues, rather than staying flat
  • captcha challenges appearing more often as concurrency increases, even against the same URL set
  • proxies getting flagged or dying faster than your pool can refresh or replace them
  • response latency creeping up, which can be the target’s own systems deliberately slowing suspicious traffic rather than an unrelated network issue

Any one of these on its own can have another explanation. All of them trending the same direction as you raise concurrency is a strong signal you’ve gone past what the pool and the target can absorb together.

Signals you’re being too conservative

The opposite failure is just as real and wastes proxy budget. If your pool is sitting mostly idle, if rented residential or mobile sessions are expiring unused before the job gets to them, or if a job is taking far longer than the target’s own rate limits would require, you’re paying for capacity you’re not using. Concurrency tuning is a two-sided problem, not just a ceiling to avoid.

Building adaptive limits instead of a fixed number

Because the safe ceiling shifts with target load, time of day, and pool health, a fixed concurrency number set once tends to age badly. Tools built for this treat concurrency as adaptive: Scrapy’s AutoThrottle extension adjusts delay based on response latency, aiohttp and httpx let you cap concurrent connections per host with a semaphore, and Playwright lets you bound how many browser contexts run in parallel. Wiring your own backoff, cutting concurrency automatically when error rates cross a threshold and easing it back up when the run is clean, does more for stability than picking a single number and hoping it holds for the whole job.

There’s no universal number

Anyone who gives you a flat “run at X concurrency” answer is guessing on your behalf. The right number depends on your pool size, the proxy type, the target’s own traffic patterns, and how its detection systems are tuned that week, and none of that is static. No proxy type and no concurrency setting makes a scraper undetectable or guarantees it won’t get blocked. Sizing concurrency to what your pool can actually sustain, and watching the error signals as you go, is what keeps a job stable. It isn’t a way to guarantee it stays invisible.

If you’re comparing proxy pools for a job like this, start with the questions providers often don’t answer upfront: pool depth, IP churn rate, and how sessions are managed. That’s the honest starting point, more useful than any single concurrency number. For more on picking proxies and running scrapers cleanly, head back to the homepage.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →