Costing a job per thousand pages: what proxy scraping actually costs
The metric everyone quotes is the wrong one
Ask a proxy vendor what their service costs and you’ll get a price per GB, a price per IP, or a price per port. None of those numbers tell you what a scraping job actually costs to run, because none of them account for the thing you’re actually trying to produce: a clean, usable page of data.
Cost per thousand pages is the metric that matters, because it’s the one your budget is actually built around. If a job needs 50,000 product pages a week, you don’t care what a gigabyte costs in isolation. You care what it costs to land 50,000 pages that parse correctly, with as little waste as possible. Everything else, the proxy tier, the rotation strategy, the retry logic, is just an input to that number.
What actually goes into the price of a page
A page has a cost even before you think about proxies: the bandwidth to fetch it, the compute to render or parse it, and the engineering time to keep the scraper working when the target changes its markup. Proxies are usually the line item people obsess over because it’s the one that scales linearly and shows up on an invoice, but it’s rarely the only thing driving the real cost.
The proxy piece itself breaks down a few ways depending on how the provider bills:
- Bandwidth-billed pools (common for residential and mobile) charge per GB moved through the network. A heavier page, more images, more JS payload, more redirects, costs more even if the proxy itself is priced the same.
- Port or IP-billed pools (common for datacenter) charge a flat rate for access regardless of how much you push through it. Bandwidth is closer to free, but the IPs are cheaper to detect and block as a class, so the real cost shows up elsewhere.
- Per-request billing exists on some managed scraping APIs, where the “proxy cost” and “request cost” are bundled into one number and JavaScript rendering or CAPTCHA handling gets billed as an add-on.
None of these are inherently cheaper. They’re just different places for the cost to hide.
The retry tax
The number that quietly wrecks a lot of cost-per-page math is the retry rate. A failed request still burns bandwidth on a metered plan, still occupies a worker, and still needs another attempt to actually land the page. If a job has a 20% failure rate on the first pass, the real cost per successful page isn’t the sticker price of one request, it’s closer to 1.25x that, once you count the wasted attempts and the second pass needed to fill the gap.
This is why a “cheap” IP pool with a high block rate against a given target can end up more expensive per successful page than a pricier pool that lands cleanly on the first try. The unit that matters is cost per successful page, not cost per request sent.
Why blocks are the real cost driver, not the proxy price
Most of the failure rate in a scraping job traces back to how the target site classifies traffic, not to the sticker price of the IP. Sites lean on a mix of signals: how consistent the request pattern looks (timing, headers, TLS fingerprint), how much history an IP or IP range has (a residential address with years of ordinary browsing behind it reads very differently than a datacenter block that’s never done anything but automated requests), and how the traffic behaves once it lands (does it request assets a browser would, does it respect rate limits, does it ever look like a human pausing to read a page).
A compliant scraper’s job is to look like ordinary, well-behaved traffic and to stay well inside a target’s published limits, not to find a clever way around its defenses. That means rate limiting to something the site can absorb, respecting robots.txt and any stated crawl policy, backing off when you see elevated error rates instead of hammering through them, and only requesting pages you actually need instead of crawling indiscriminately. None of this guarantees you won’t get rate limited or blocked. Sites tune their detection constantly, and no proxy type or vendor can promise otherwise. What good practice does is keep your job on the right side of a site’s terms and reduce unnecessary waste, which is also, not coincidentally, what keeps your cost per page down.
A framework for working out cost per thousand pages
The formula is simple; the discipline is in tracking the inputs honestly:
Cost per 1,000 pages = (proxy spend + retry overhead + tooling/compute spend) ÷ successful pages × 1,000
To make this concrete, walk through the pieces with round numbers used purely to illustrate the math, not as a claim about what any provider actually charges:
- Say your average page, with its assets, pulls a certain amount of data per request.
- Multiply that by your total request volume, including retries, to get total bandwidth for the job.
- Apply your provider’s per-GB or per-port rate to get raw proxy spend.
- Divide by the number of pages that actually came back usable, not the number of requests sent.
- Add in whatever you’re spending on rendering (headless browser compute is not free), parsing maintenance, and any CAPTCHA-solving or verification step the job requires.
The output is a real, comparable number. It’s also one that changes over time, because failure rates against a given target drift as that target adjusts its defenses. A job that cost X per thousand pages last quarter can quietly cost more this quarter with zero change on your end, simply because the target got better at spotting the pattern you were using. Tracking this number on a recurring basis, not just at project kickoff, is the only way to catch that drift before it eats a budget.
Datacenter vs residential vs mobile, priced by outcome
The honest answer to “which is cheapest per page” is that it depends entirely on the target, not on the proxy type in isolation.
Datacenter IPs are cheap per port and often unmetered on bandwidth, which makes them attractive on paper. Against a target that doesn’t scrutinize IP reputation closely, they can produce a very low cost per page. Against a target that filters datacenter ASNs aggressively, the failure rate can be high enough that the effective cost per successful page is worse than a pricier pool, once retries are counted.
Residential IPs cost more per GB because you’re paying for address diversity and a reputation that looks like ordinary consumer traffic. Against targets with strong anti-bot tooling, that reputation often translates into a meaningfully lower failure rate, which can bring the real cost per successful page down even though the sticker price per GB is higher.
Mobile IPs sit a step further in the same direction: carrier-grade NAT means many users share the same address, which changes how a site can even attribute suspicious behavior to a single source. That can help against targets tuned specifically for residential detection, but it comes at the highest bandwidth cost of the three, so it only pays off if the failure-rate improvement is large enough to offset it.
There’s no universal ranking here. The only way to know which tier is cheapest for a specific job is to run a small, honest test against the actual target, track successful pages against total spend, and compare the resulting cost per thousand across tiers before committing a full job to one.
Where the real savings are
Most of the durable savings in a scraping job come from reducing waste, not from finding a cheaper proxy. Caching pages you don’t need to refetch, blocking unnecessary asset loads when you only need the HTML, narrowing the crawl to the specific pages a job requires instead of crawling broadly, and monitoring success rate closely enough to catch a target’s changes early all do more for cost per page than shopping for a marginally cheaper GB rate.
Proxies are one input to a job’s economics. Treating them as the whole budget is how teams end up optimizing the wrong number.
If you want more breakdowns like this on running proxy-based scraping honestly and at scale, the rest of what we cover is on the Proxy Scraping homepage.
Get new guides and videos first — join the Telegram channel.