Choosing a concurrency limit from the target's own signals
The number you picked on day one is already wrong
Most scrapers start with a concurrency limit that came from nowhere real. Someone picks 20 threads, or 50, or “as many as the proxy pool can support,” and ships it. That number reflects what the scraper’s own infrastructure can handle, not what the target site can handle. Those are two different ceilings, and the second one is the one that actually determines whether requests come back clean or come back as errors, CAPTCHAs, or silence.
A target’s own signals tell you where its ceiling is. Latency, status codes, headers, and connection behavior are all data the site is already handing back on every request. Ignoring them and running a static concurrency number is like driving with your eyes closed and trusting the speedometer alone. The road is telling you something. A scraper that reads it runs longer, gets banned less, and produces cleaner data.
What “the target’s signals” actually means
Every response carries more than the payload. It carries timing, a status code, sometimes explicit rate limit headers, and sometimes the shape of a failure (timeout vs. reset vs. a served block page). Together these are the target’s side of a feedback loop. A concurrency limit that’s tuned to this feedback is not a fixed number, it’s a moving one that tracks how much load the site is currently willing to absorb.
This matters because sites don’t publish their true capacity. A site might handle 200 concurrent requests from a residential IP range fine at 3am and start serving degraded responses at the same load during a traffic spike at 2pm. The only way to know the ceiling in real time is to watch what comes back.
Response latency drift
The clearest early signal is latency creeping up under load that used to be fine. If the median response time for a given endpoint doubles or triples as concurrency rises, that’s the target’s infrastructure (or a load balancer, or a WAF doing more inspection) working harder to serve the traffic. Latency drift usually arrives before outright blocking does. A scraper that tracks p50 and p95 response times per target and backs off when they climb past a baseline catches trouble before it becomes an error rate problem.
The practical version of this: keep a rolling window of response times, say the last 100 requests to a given host. If the p95 crosses some multiple of its steady-state value, cut concurrency for that host rather than pushing harder. Pushing harder into rising latency almost always makes it worse, because queued connections pile up on both ends.
Status codes and what they’re actually telling you
HTTP status codes are the most direct signal a target sends, and they mean specific things:
- 429 (Too Many Requests) is explicit. The site is telling the client it’s over the allowed rate. This is not ambiguous and should always trigger a backoff, never a retry at the same rate.
- 503 (Service Unavailable), especially when it starts appearing under load and disappears at lower concurrency, usually means the origin or a proxy in front of it is shedding load. Some 503s come with a
Retry-Afterheader; that header is the target literally stating how long to wait, and it should be honored, not treated as a suggestion. - 403 (Forbidden) appearing intermittently, correlated with request volume from a given IP or session, usually means a WAF or bot-detection layer has started challenging or blocking that source. A rising 403 rate at a given concurrency is a hard ceiling signal, not something to push through.
- CAPTCHA or interstitial challenge pages served instead of the expected content are a site protecting itself from load it considers abnormal. From an operator’s side, this is a stop signal: the site has decided the current pattern looks automated or excessive, and continuing to hammer it does not change that decision, it just accumulates more blocked sessions and burned IPs.
None of these are obstacles to route around. They are the target’s own rate control system, and reading them correctly is what lets a scraper adjust before it accumulates a pile of wasted, blocked requests.
Connection-level signals
Below the HTTP layer, TCP resets, connection refusals, and timeouts carry information too. A site that starts silently dropping connections instead of returning any HTTP status is often doing coarser-grained blocking, sometimes at the network edge rather than the application layer. This tends to show up when concurrency from a single IP or a narrow IP range gets high enough to look like a single aggressive client rather than organic traffic. Tracking connection failure rate separately from HTTP error rate catches this, because a scraper that only checks response.status_code never sees a connection that never completed.
Reading target signals alongside proxy pool signals
Concurrency limits do not exist in isolation from the proxy layer. The same total request volume looks completely different to a target depending on how it is distributed. A hundred requests a minute from one datacenter IP looks like a single very active client. The same hundred requests spread across a rotating pool of residential or mobile IPs looks like a hundred different low-volume users, because it is coming through carrier-grade or ISP-assigned addresses that naturally see irregular traffic.
This is why concurrency tuning and proxy pool sizing are the same conversation, not two separate ones. If a target’s error and challenge rate rises at a given concurrency, the fix is not always “reduce concurrency,” it can also be “reduce concurrency per IP” by widening the pool, since the signal the target is reacting to is often per-source load rather than total load. Datacenter ranges are the easiest for a site to fingerprint and rate-limit as a block, because they’re well-documented ASN ranges. Residential and mobile ranges look like ordinary consumer traffic, which changes what “concurrency” even means at the target: five requests a minute from each of twenty residential IPs reads very differently to a rate limiter than a hundred requests a minute from one IP, even though the total volume is identical.
None of this changes what the signals mean. A 429 is still a 429 regardless of what kind of proxy sent the request that triggered it. It just changes where the fix gets applied: sometimes it’s total concurrency, sometimes it’s concurrency per source, and the target’s error rate at each configuration is the only reliable way to tell which one is the actual constraint.
A practical way to set the ceiling
Start low. Run a small, fixed concurrency against a target and record baseline latency and error rate. Step concurrency up gradually, watching the same two numbers. There is usually an inflection point where latency and error rate stay flat for a while and then start rising sharply rather than gradually. That inflection point, not the highest number that technically still returns 200s, is the real ceiling. Running right at the inflection point means the next small spike in the target’s own load, unrelated to the scraper at all, pushes into block territory.
A sensible operating concurrency sits comfortably below that inflection point, with an automated backoff that kicks in on any of: 429 responses, a rising 503 or 403 rate, latency crossing a threshold above baseline, or a rise in connection failures. The backoff should reduce concurrency proportionally, not just pause and resume at the same rate, since resuming at the same rate that triggered the block just triggers it again.
When the answer is to stop, not slow down
Some signals mean the target does not want this traffic pattern at any concurrency, and the correct response is to stop rather than tune. Repeated CAPTCHA challenges after backoff, a site’s robots.txt disallowing the path being requested, or terms of service that prohibit automated access are not rate-limiting problems to solve with better pacing. Concurrency tuning is for staying within a target’s technical capacity gracefully. It’s not a way to get around a target’s stated policy on being scraped at all. Those are different questions, and no amount of signal-reading changes the answer to the second one.
Concurrency that respects a target’s own feedback produces better data anyway: fewer retries, fewer blocked sessions to recover from, and a scraper that keeps running next week instead of getting an entire IP range flagged this week.
Read more on proxy selection, rotation, and honest provider comparisons at the Proxy Scraping home page.
Get new guides and videos first — join the Telegram channel.