← all guides

Handling retries without amplifying your own block rate

Most scrapers don’t get blocked because a target caught something clever. They get blocked because their retry logic panicked and hammered the same endpoint five times in two seconds after the first request failed. The block wasn’t inevitable. It was self-inflicted, and the retry strategy caused it.

I’ve run proxy infrastructure for scraping operations long enough to see this pattern over and over: a job that was working fine at a 2% error rate suddenly spikes to 40% blocked, and the root cause traces back to a retry loop that treated every failure the same way. This piece is about the mechanics of why that happens and how to build retry handling that doesn’t compound the problem it’s trying to solve.

Why a failed request isn’t one signal, it’s several

The first mistake is treating “the request failed” as a single category. In practice a failed request from a scraping target can mean several different things, and each one calls for a different response:

  • A 429 with a Retry-After header means the target’s rate limiter is telling you exactly how long to wait. Ignoring that header and retrying immediately is the single fastest way to escalate a soft rate limit into a hard block.
  • A 403 or a redirect to a challenge page usually means something in the request (headers, TLS fingerprint, IP reputation, behavioral pattern) tripped a detection layer, not a capacity limit. Retrying the same request from the same IP just confirms the pattern that got flagged in the first place.
  • A 503 or connection timeout can be genuine infrastructure load on the target’s side, unrelated to your traffic at all.
  • A CAPTCHA wall is a signal that the target wants a human in the loop. Automating past it isn’t a retry problem, it’s a decision about whether that data source is appropriate for automated collection at all.

A retry strategy that fires the same backoff-and-retry logic regardless of which of these happened is treating a rate limiter, a bot detector, and a load balancer as the same system. They aren’t. Reading the actual response, status code, and headers before deciding whether and how to retry is the first piece of discipline that separates a clean scraper from a noisy one.

The thundering herd problem, at the scale of one job

“Thundering herd” usually gets discussed at the scale of distributed systems, but it shows up just as badly inside a single scraping job. If you’re running concurrent workers against the same domain and a shared rate limit or block trips, naive retry logic means all of them fail around the same time and all of them retry around the same time. Now you’ve turned one transient block into a synchronized burst that looks, from the target’s side, exactly like an automated attack pattern: a wall of requests arriving in a tight window right after a rejection.

This is where fixed-interval retries do real damage. “Wait 2 seconds and try again” sounds reasonable for one worker. Multiplied across fifty concurrent workers that all failed at once, it produces fifty requests landing in the same one-second window, twice. Any reasonably built anomaly detector treats a burst like that as confirmation, not coincidence.

Jitter fixes this specifically. Instead of every worker retrying at exactly 2 seconds, each one retries at a randomized point in a window, say 1 to 4 seconds. That single change breaks the synchronization and spreads retries out enough that they look like independent, uncorrelated requests instead of a coordinated burst. It costs almost nothing to implement and it’s one of the highest-leverage changes you can make to a retry system.

Exponential backoff is necessary, but not sufficient

Exponential backoff, doubling the wait time after each consecutive failure, is standard advice and it’s correct advice. But it only solves the timing problem for a single failure chain. It doesn’t solve the identity problem.

If a worker gets a 403 on attempt one, waits, retries, gets another 403, waits longer, retries again, and gets a third 403, all from the same IP and the same fingerprint, exponential backoff has made the retries slower but it hasn’t changed anything about why they’re failing. At some point continuing to retry from an identity that’s already been flagged isn’t patience, it’s just generating more data points that confirm the flag. A retry strategy needs a ceiling: a maximum number of attempts against a given target from a given identity before the system stops and escalates the failure to a human or a different handling path, rather than retrying indefinitely with longer and longer gaps.

Retries and rotation have to be coordinated, not independent

Rotating proxies is often treated as the fix for blocked retries: request fails, rotate to a new IP, retry. That can work, but only if the rotation and the retry logic are actually talking to each other. Two failure modes show up constantly when they aren’t:

The first is rotating on every single retry regardless of cause. If a request failed because of a 503 (target-side load, nothing to do with your IP), rotating proxies and retrying doesn’t address the actual problem and just burns through proxy inventory for no reason. Rotation should be a response to identity-specific signals (403s, CAPTCHA walls, IP-level rate limits), not a blanket reflex for every failure type.

The second, and more damaging, is rotating fast enough that the request pattern itself becomes the tell. A target that sees the same session cookie or the same request fingerprint arrive from five different IPs inside thirty seconds isn’t looking at five independent visitors, it’s looking at one scraper cycling through proxies, which is arguably a stronger signal than a single IP retrying too fast. Rotation needs pacing behind it, not just a fresh IP on every attempt.

Tracking your own block rate as a metric, not an afterthought

The operators who keep clean retry behavior are the ones who watch their own block rate as a live number, not something they notice after a job finishes at 60% failure. That means logging, per target domain, the ratio of blocked or challenged responses to total requests over a rolling window, and having retry logic that reads that number back.

A practical version of this: if the rolling block rate against a domain crosses some threshold, say 15% over the last hundred requests, the system should slow down globally for that domain, not just retry the individual failed request faster. That’s a circuit breaker pattern, and it exists specifically because per-request retry logic has no visibility into the aggregate pattern it’s creating. A single request doesn’t know it’s the fortieth retry against the same domain in the last minute. Something above it needs to know that, and needs the authority to pause the whole job rather than let every worker keep independently retrying into a target that’s already showing signs of tightening up.

This is also the point where a scraper should be willing to just stop. Not every blocked target is worth pushing through with more retries, more rotation, or more patience. If a site is clearly enforcing access controls around content that requires authentication, sits behind a paywall, or is otherwise signaling it doesn’t want automated collection, escalating retry attempts against that signal isn’t a technical problem to solve, it’s a decision about whether the job should be running at all.

What this looks like put together

A retry strategy that doesn’t amplify block rate has a few concrete properties: it reads the actual failure type instead of treating every non-200 the same, it respects Retry-After when a target provides one, it uses jittered backoff instead of fixed intervals so concurrent workers don’t retry in sync, it caps attempts per identity instead of retrying forever, it coordinates rotation with retry cause instead of rotating reflexively, and it tracks aggregate block rate per target so the system can slow down or stop before one blocked request becomes a hundred. None of these are exotic techniques. They’re closer to basic distributed systems hygiene, applied to a context where the “other side” is actively watching for patterns that look automated, and where getting the pattern wrong doesn’t just cost you a retry, it costs you the IP, the session, and sometimes the account.

If you’re building or debugging scraping infrastructure and want more of this kind of grounded, defensive-side explanation of how proxies, rotation, and blocking actually work, you can find the rest of our writing and videos at Proxy Scraping.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →