← all guides

When a proxy is the wrong fix

proxies web-scraping bot-detection buying-proxies diagnostics

A job that dies at request 800 from one address dies at request 8,000 from ten. The wall moved by exactly the factor you paid for. It is still there.

That arithmetic is most of what I have to say, and I sell the addresses, so take the rest with that in mind.

Real SIM cards on Singapore carriers, modems in a rack in my flat, roughly $10 a month in data per line and another $1.50 in modem depreciation. An order for forty lines instead of four is a real difference to my month, and I keep turning those orders down because the follow up refund is worse than the sale.

The check that settles it

Open a normal browser. Load the exact URL your scraper was asking for when it failed. Do it from the same address your scraper uses.

If the page comes back, the address is clean. The site is willing to serve that content to that address, right now, to somebody. Whatever refused your code was something your code did.

While you are in there, open the network tab. Every request the page fires is listed with its headers and its ordering. You are looking at a free reference copy of what a legitimate visit to that URL looks like, which is the shape your client is failing to match.

The way almost everyone runs it wrong

They open the browser on a laptop at home and run the scraper on a VPS in Frankfurt.

Two different addresses, so the check proved nothing. If the job runs on a server, the request has to leave from that server: SSH in and fetch it there, or route a browser through the same exit your code uses. That is the step people skip, and skipping it is how they talk themselves into a purchase.

One more way it lies.

If the page loads and the contents are quietly wrong, prices served from a stale cache, a listing truncated at twenty rows when you know there are 400, you have passed nothing. Confirm that what came back is what you wanted, not just that something came back.

So what did it, then

Assume the page loaded. The address is fine. Two of the causes are well worn, so I will be brief about them.

Pace, and the total absence of think time. A person opens a page and then does nothing for forty seconds because they are reading. Your code has no reading step, so the gap between requests is whatever the network cost you.

Adding time.sleep(2) produces a gap that is identical every single time, and a constant gap is easier to count than a fast one.

Headers. Your client says it is Chrome 131 and sends four of them. A browser sends twelve to eighteen in a stable order, including Accept-Language and the fetch metadata describing how the page is being loaded. The request contradicts itself, and catching a contradiction takes an if statement.

Now the three people miss.

The order you ask in

A person reaches a site through the front or through a search result, and then clicks. Category, product, back, another product, a link buried in a description.

The path is a walk. It is messy and it always starts somewhere reachable.

Your scraper opens page 400 of a listing with no referrer, having never seen page 1. Then 401. Then 402. It never goes back, never opens the same item twice the way people constantly do, never strays into a section it does not need.

Spotting that needs no address reputation and no model. It is a sequence, and it is a sequence no human generates.

I watched a job get flagged inside 200 requests on nothing but this. Mobile addresses, a polite two second cadence, headers from a real browser because it was a real browser. It walked the catalogue in perfect numerical order like a librarian with no other interests.

The endpoints that exist to be watched

Some URLs are instrumented far harder than others, and the hardest watched are the ones only worth hitting if you are automating.

The JSON feed sitting behind the product page. The internal search API the front end calls. The ?limit=100 that returns five pages worth of rows in one response when the interface shows twenty. The bulk lookup nobody has ever clicked because no button points at it.

Those are the convenient ones, which is precisely why they are counted. A site that has been scraped before knows exactly which of its own paths scrapers love, because it has watched them get hammered for years.

So you can be gentle, correctly dressed, sitting on a clean carrier address, and still be the only party in the log requesting that path at that volume. The address never entered the decision. You selected yourself.

Logged in work, where no address helps

If the job authenticates, the account is the identity and the address is a detail hanging off it. Rotate through 400 exits and every request still arrives carrying a cookie that names you.

It is worse than neutral. An account whose session hops across forty addresses in three cities inside ten minutes reads as stranger than one that sat still. You have taken the signal that made the session boring and made it interesting.

People buy pools specifically to protect logged in work, and it is the clearest case of buying the wrong thing. The account carries the risk. Spend the money on however many accounts you are genuinely permitted to hold, and give each one a stable address it keeps for its whole life.

What the pool costs you besides money

Here is the part I learned the expensive way.

A large pool makes the failure harder to read.

On one address, everything that happens to you happens in one place and accumulates. You can watch a run degrade. You can find the request where the tone changed.

Spread the same job over forty addresses and each carries a fortieth of the story. No single line gets far enough into the pattern to show you the pattern. What was a cliff in your error rate becomes a gentle fog at 2%, and 2% of everything is far harder to chase than one address dying at request 800.

A customer once sent me forty lines worth of logs and asked which address was the bad one. None were. The job failed at the same rate on every line, evenly, which is the signature of a behaviour problem and the hardest shape to see once it has been smeared that thin.

So the first purchase avoids the diagnosis and then makes the next one harder. That is the position I will defend: most people’s first proxy order is a way of avoiding their own traffic. It costs money instead of attention, and money is by far the cheaper of those to spend at 11pm on a job that has been failing since lunch.

When it really is the address

Sometimes it genuinely is, and leaving that out would be its own kind of lie. Five cases where I take the order without arguing.

You are on a datacenter block. A cloud IP gets refused before behaviour is considered, because the range belongs to a hosting company, the ASN is published, and some sites drop the whole thing on sight. This is the one case where the browser check disagrees with itself: the page loads from your home line and dies from the server. That gap is your answer.

The content is geographic. If a site serves different data by country or refuses you outside it, no amount of good manners produces the right page. You have to be somewhere else, and being somewhere else is a purchase.

You inherited someone’s mess. On a cheap shared pool the address may have been burned by whoever held it last week. Worth swapping, though worth swapping to a better supplier rather than to more of the same pool.

The volume is legitimate and genuinely large. One address cannot plausibly request two million pages, because no person does. That is a case for parallel identities chosen deliberately, each behaving like a separate visitor with its own history and its own session. Nothing like the panic order placed twenty minutes after a block.

Blast radius. Buying a separate line so the address that eventually gets burned is one you can afford to lose, while the office connection everyone else works over stays clean. That is insurance. Sensible, and it repairs nothing.

The question I ask before quoting

Somebody asks for a bigger pool. I ask what changed between the first request and the one that failed.

Not the address. The behaviour. Did it fail at a consistent count or at a consistent hour. Did the error land on one URL or across all of them. Did the response body change before the status code did.

If they can answer, I quote, and it is usually a smaller order than they walked in for. If the answer is that they do not know, more addresses buy a delay and cost a diagnosis.

What I cannot tell you

Which of these is yours. I see requests and responses from outside and I infer, and I have been wrong about it.

I also cannot promise the browser check never misleads. A site can decide to serve a person and refuse a program from the same address, on purpose, and then the address really is clean and the work is still impossible. That happens on the large platforms, and my advice there has not moved: buy the data, use the official API, or drop the project. The test in the form I hand customers, and the numbers I ask them to record before they order anything, are here.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →