← all guides

Where you run your scraper matters

web-scraping scraper-hosting latency dns-leaks scraping-infrastructure

Someone wrote to me in May to say my lines had gone slow. Same code, same pool, and his nightly job had gone from finishing before breakfast to still running at noon.

Nothing had changed on my side. Over the weekend he had moved the scraper off a desktop in Singapore onto a rented server in Ohio, because the server was four dollars a month and the desktop was a desktop.

My SIM cards sit on a shelf in my flat. So every request now crossed the Pacific and came back before the proxy had done any work at all.

About 220 milliseconds of it, on every request, before the request had got anywhere.

What the proxy actually decides

The exit owns the identity. That much is correct and I will not argue with it.

The address, the network it belongs to, the country it geolocates to, whether it reads as a phone or a rack. All of that comes from where the traffic left. You can run the scraper on a Raspberry Pi under a desk in Singapore and the site will file you as a residential connection in Frankfurt.

Which makes the next step easy. If the target cannot see my machine, my machine is not part of the problem.

Reasonable enough. It holds up for about a week.

The first hop is on your bill

Your traffic does not materialise inside the proxy. It travels there first, and that leg belongs to you.

It also hides from the one test everybody runs. You point a speed checker through the pool, you get 300ms, you shrug, because 300ms is nothing on a page load.

Then you send 400,000 requests. 220 milliseconds of avoidable distance across 400,000 requests is 24 hours of a process sitting still while packets walk.

It never presents as 24 hours, because concurrency eats it. At 16 in flight that compresses to about 90 minutes and the run time looks healthy. What it cost was the 16. Nine would have moved the same volume from the right region.

And how many you have in flight is the one number your target is genuinely counting, which is an argument I have made at length elsewhere. Short version: you want it as low as your deadline tolerates.

So the cheap box did not cost him hours. It cost him concurrency, and concurrency is the expensive thing.

This is the one part of the decision you can settle with data. Rent the cheapest instance in three candidate regions, time 500 requests through your real pool at your real target, keep the winner and kill the rest. An hour of work, well under a dollar.

The machine has to still be there at 4am

A six hour job needs a machine that stays awake for six hours, and a laptop is not that machine.

The lid closes and it sleeps. Or nothing touches it for twenty minutes and it sleeps. Or the battery drops past twenty percent and something sensible takes over on your behalf. I came back to an overnight run once showing hour two. The lid had been down for seven hours and the process had been suspended that entire time, politely, exactly as designed.

Scheduled work fails worse than that. A cron entry on a sleeping machine does not queue up and fire late. The hour passes, nothing runs, and your log has no line in it for you to be suspicious about.

Home connections do their own version. Router firmware at 4am, wifi roaming to the other access point mid job, an update that reboots the machine and restores your browser tabs but not the python process doing the work.

None of it arrives as a clean stop either. Sockets hang, half the workers park in a read that will never return, and the job neither finishes nor fails. Without a checkpoint you start again from zero, which means a second full pass at a target that has only just finished watching the first one.

Your resolver is a second network

The proxy carries your requests. It does not automatically carry your name lookups.

With an HTTP proxy and a CONNECT tunnel the proxy resolves the hostname for you. With SOCKS you get a choice, and the two spellings are one letter apart. socks5 resolves locally and hands the proxy an address. socks5h hands the proxy the name.

Pick the wrong one and every hostname you scrape goes to your own resolver first, from your own address, unencrypted.

There are other doors into the same room. A headless browser configured in a hurry. A monitoring check that pings the host by name from the box rather than through the pool.

What leaks is the hostname, and only the hostname. But the hostname is the entire point. A resolver log showing one connection asking for the same fifty names at 03:00 every night is not an ambiguous document.

The answers move too. A resolver in one country hands back a different edge address than a resolver in another, so you exit in Germany and arrive at a node in Asia. That matters far more once your target sits behind a CDN.

Test it rather than assume. Send one request through the pool to something that reports back which resolver it saw. If the answer names your own network you have found it in a minute.

When the wrong time reads as a wall

I lost a day and a half to this one.

Every HTTPS request failing at once, through every proxy in the pool. Certificate errors read as interference when they arrive through a proxy, because your first instinct is that something has sat down in the middle of the connection.

The clock on the box was three days behind. It was a VM restored from a snapshot and nothing had corrected it since.

A certificate issued yesterday is not yet valid on a machine that believes it is last Tuesday. Every handshake fails, the client library reports a plain connection error, and off you go hunting for a block that never existed.

The fix is boring. One time sync service, running, verified, and confirm the container inherited it rather than assuming. Cheap servers ship without one more often than you would expect, and anything signing a timed token breaks in the same shape on a few minutes of skew.

Development stays where you are

I want to be clear about this, because the internet’s answer to every question is to rent something.

Writing the scraper. Running it while you watch it. A one off pull of 4,000 pages while you eat lunch. All of that belongs on the machine already in front of you, and your home connection is often the shortest path to the proxy anyway, because it is near you and a cheap server in somebody else’s country is not.

The laptop earns its place. It earns it for work with a person attached.

Anything scheduled gets a box of its own

The test is whether somebody is present when the job runs.

At 3am nobody is present, by definition, and the lid becomes a single point of failure that a tired human closes every night.

Small is enough. One core and a gigabyte of memory will run a request based scraper at a rate your target objects to long before the box does. Headless browsers change that arithmetic, mostly on memory, and further than most people expect.

Rent near the proxy

Here is the part I had backwards for years, and most people I talk to have it backwards too.

The instinct is to sit close to the site you are scraping. Shorter path to the data. It feels obviously correct.

But you have no path to the site. You have a path to the proxy, the proxy has a path to the site, and only the first one is yours to pick. The proxy leg happens on every request wherever you sit, so shorten the leg you own and stop thinking about the other one.

Mobile lines in Singapore, rent in Singapore. Exits in Europe, rent in Europe. The target can be wherever it likes, because the exit already decided what it sees.

And if you authenticate by whitelist, update the list before you move the box rather than after, for reasons I have written about elsewhere.

Five dollars was never the obstacle

A small box is a few dollars a month. I pay ten for a single SIM card and another 1.50 in modem depreciation, and the SIM is the part that makes any of this work at all.

Nobody keeps a scheduled scraper on a laptop to save that. They keep it there because it worked the day it was written, it kept working, and there was never a morning where it broke loudly enough to force a move.

If a job runs on a schedule and it runs on a machine you also use for other things, you have not chosen where it runs. You have never moved it.

People argue with me about this and it is always the same argument. It has been fine for a year. And it has been fine for a year, right up until the row count quietly halves and you spend three days reading proxy logs about a laptop that went to sleep.

What I cannot tell you

Near the proxy assumes there is a near. A pool spread over six countries does not have one, and then you are averaging. My only answer is to sit in the region most of your traffic exits from and accept the rest, which is not much of an answer.

I also cannot tell you where a bigger machine beats a closer one. I have never run a scrape where the processor was the constraint, but I do not do heavy browser work at volume either, so treat that as a gap in my experience rather than a finding.

And I still leave real jobs on my laptop. Did it last month. The only thing that has changed is that I notice inside a day now, instead of finding out from a row count three weeks later.

The latency check I run before picking a region, and the list I work through when I set up a new scraper box, are here.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →