← all guides

Why your scraper gets blocked (and what actually fixes it)

web-scraping proxies bot-detection rate-limiting

A job of mine used to die at around request 800. Every run, give or take a few dozen. I decided the exit pool was too thin, moved the whole thing onto mobile lines, and waited for the win.

It died at request 900.

That was the entire return on three weeks and a hardware bill. The real cause was sitting in my own retry handler: on any error it retried immediately, three times, against the same URL. Every hiccup on their end turned into four rapid hits from me, which produced another error, which produced another four. My scraper was manufacturing the burst that got it counted.

I sell mobile proxies. I sold myself the wrong thing.

The four causes, in the order that matters

  1. How fast you ask, and how regular the gaps are
  2. Whether your request looks like the software it claims to be
  3. Whether you hold a session or throw one away every time
  4. The reputation of the address you came from

Almost everybody starts at four. It’s the one you can fix with a card payment, which is most of the explanation. It also has a sales team attached, and I’m on it, so read the rest accordingly.

Rate, and the thing people miss about it

Rate limiting predates every fancy detection system and still catches more scrapers than all of them combined. No site has to prove you’re a bot to act. It only has to count: one address asked for 900 pages in four minutes, the median visitor asks for six.

That check costs nothing to run and it’s usually the first rule anyone writes. Which is why the first fix is boring. Slow down. I’ve taken a job from a block every few hundred requests to a clean run of tens of thousands by editing one number.

But the number most people edit is concurrency, and the number that actually gives them away is variance.

Two requests a second, forever, exactly on the second, is more conspicuous than five requests in one second followed by nineteen seconds of silence. Human traffic is lumpy. Someone opens a page, reads for forty seconds, fires three tabs at once, then goes to make coffee.

A fixed time.sleep(2) produces a metronome. The mean looks reasonable and the standard deviation is zero, and zero is a fingerprint anyone can compute in SQL.

So randomise the gap. Randomise the burst size. Stop the run completely sometimes, because people sleep and cron doesn’t. And stagger your start off the hour, since every scheduled job on the internet fires at :00 and that’s the most watched minute in any log.

When your client contradicts itself

Second cause: your request says one thing about itself in one place and something different in another.

The classic is a user agent set to the newest Chrome build, sent from a Python library carrying four headers. A real browser sends twelve to eighteen, in a stable order, including Accept-Language and the fetch metadata headers describing how the page is being loaded. So the request announces itself as a browser while carrying almost none of a browser’s luggage. Catching that takes an if statement.

It goes below the headers too. The TLS handshake your client opens with has a shape of its own, from the cipher list and the ordering of extensions, and that differs between a real browser and a scripting library. A server sees it before a single header.

I’m not going to walk you through forging any of that, and forging it is the wrong repair anyway. Remove the contradiction instead. Two honest ways:

  • Drive an actual browser, so the headers and the handshake belong to a browser because one genuinely produced them
  • Send your own identifier, name yourself, put a contact URL in it, and read the site’s robots file and terms first so what you collect is what you’re permitted to collect

The second sounds hopelessly naive. It works more often than people expect on smaller sites, and some operators will hand over the data or point you at an API if you ask.

The session you keep throwing away

Third cause, and the one that quietly wrecks the most jobs.

A normal visitor lands, picks up a cookie, and carries it around. Their next twenty requests are visibly the same person continuing. A scraper with no cookie jar, or a fresh jar per request, is a thousand separate first-time visitors from the same neighbourhood, each with no history, each going straight to a deep URL nobody reaches without clicking something first.

Spotting that needs no address reputation. It’s pure sequence.

So keep the jar. Let the session accumulate. Walk in the front door occasionally instead of teleporting to page 400 of a listing.

Why rotating on every request is usually worse

The default in nearly every proxy tutorial is a fresh exit on each request. It feels cautious. Spread the load, no single address does much, nothing to count.

Against a site with any session tracking, it’s worse than doing nothing. You now have one continuous session identity hopping between 200 addresses in different cities inside a couple of minutes. No person browses like that. A bank would freeze your card over it. And you’ve discarded the history that made the session look ordinary.

Hold one address for the length of a session, then rotate between sessions. That matches how people actually browse, and it’s less code.

Per-request rotation has real uses: independent single-page fetches with no state between them, price checks, availability polls. If your job has a login, a basket, or a search that carries filters across several steps, holding wins and it isn’t close.

The address, last

Now the part I sell.

A datacenter address gets caught on reputation alone, before behaviour is even considered. The block belongs to a hosting company, the network number is public, and anyone can see no household has ever lived there. Some sites refuse the whole range. Your request was rejected in the second it arrived.

A residential or mobile address usually clears that filter, because it sits in a block a carrier hands out to ordinary phones and homes, and refusing the block takes paying customers down with it. A mobile address in particular sits behind carrier-grade address translation with a crowd of other subscribers, so it’s hard to pin to one person.

All of that is reputation. It’s a filter you pass, and anybody telling you a proxy type can’t be blocked is selling. Behaviour still catches you on any address. I’ve watched a mobile line get blocked in about ninety seconds because the scraper behind it was doing something absurd.

What a good address buys is that the cheap first filter never fires, so your behaviour becomes the thing being judged. That’s worth paying for. It’s worth much less than fixing the behaviour, which is why I keep talking buyers down from forty lines to four. At roughly $10 a month in data per line plus about $1.50 in modem depreciation, the gap between four and forty is real money that mostly buys nothing.

The three ways a site says no

They aren’t equally dangerous, and the middle one is the one that costs you.

A hard block is a 403, a 429, or a dead connection. This one is honest. Your error rate spikes and you go and look.

A soft block returns 200. The page renders, the data is there, and it’s subtly wrong: prices served stale from a cache, a listing truncated to the first twenty rows when there were 400, a field that used to populate now empty on every record.

The soft one is expensive precisely because your monitoring stays green. The job succeeded. Row counts look normal. And you’re writing rubbish into a database and reporting on it for weeks.

A challenge page is an interstitial asking for proof before it lets you through. That’s the site saying plainly it wants a human here. Treat it as an answer. Back off, and if it persists, find the official API or ask for access. Punching through a challenge is where scraping stops being scraping.

Catching a soft block

Build a canary. Pick five records whose true values you know, verify them by hand once a week, and have the job assert against them on every run.

Then monitor the shape of the output, not only its arrival. Field fill rates. Rows per page. Response size in bytes. All of those move when a site starts serving something different. None of them move if you only check status codes.

The order to test in

Someone tells me their scraper is blocked. This is the sequence, free until the last step.

Open the same URL in a normal browser on your home connection. If that works and your scraper doesn’t, the address was never the issue and you’ve avoided a purchase.

Set concurrency to 1 and the delay to ten seconds, then run twenty requests. Clean result means you have a rate problem.

In devtools, copy the browser’s request as cURL and diff it against what your client sends. The gap is usually embarrassing.

Check whether you’re persisting cookies across requests. Plenty of people find out here that they aren’t.

Only then does buying addresses make sense, and by then you know precisely what you need them for.

What I can’t tell you

Some targets I’ve never got working politely. A few large platforms state in their terms that they don’t want automated collection, and my advice there is to buy the data, use the official API, or drop the project.

I also can’t tell you what any specific site runs. I see a request and a response and I infer, and I’m wrong sometimes. Anyone who tells you exactly which system blocked you, from the outside, is guessing. The written guides and the proxy setups I actually start from are here.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →