Costing a scrape by bandwidth instead of requests
Most people budget a scrape by counting requests. Ten thousand pages, ten thousand requests, done. That number is almost useless if you’re paying for proxies by the gigabyte, which is how residential and mobile proxy pricing usually works. Two scrapes hitting the same number of pages can cost wildly different amounts, because the request count says nothing about how many bytes came back each time.
I run proxy infrastructure and scrapers against it daily, and the bandwidth line is the one that actually blows budgets. Request count is a proxy for cost. Bandwidth is the cost. Here’s how to think about the difference and build a model that doesn’t surprise you at the end of the month.
The request count lies to you
A request to a JSON API endpoint that returns a few kilobytes of structured data costs almost nothing to pull. A request to a fully rendered e-commerce category page, with product images, tracking scripts, web fonts, and a few megabytes of JavaScript, can cost a thousand times more in transferred bytes for the exact same “one request” in your log.
If your proxy plan is billed per gigabyte, that difference is your bill. Ten thousand API calls might total a few hundred megabytes. Ten thousand rendered product pages, especially if you’re using a headless browser that loads every asset on the page instead of just the HTML, can run into tens of gigabytes. Same request count, completely different invoice.
This is why “how many pages do I need” is the wrong first question for a budget. The right first question is “how many bytes does one target page actually transfer, and how much of that do I actually need.”
Where the bytes actually go
Page weight isn’t evenly distributed. On a typical content-heavy page, the HTML itself is usually a small fraction of total transfer. The rest is images, video, web fonts, CSS, and JavaScript bundles, plus whatever third-party trackers, ad tags, and chat widgets the site has bolted on. None of that is data you’re scraping. It’s data your client downloaded on the way to rendering the data you’re scraping.
A few things multiply bandwidth beyond the page weight itself:
Redirects. A single logical request that bounces through two or three redirects before landing on the final page burns bandwidth on every hop, and each hop can carry its own tracking pixels and headers.
Retries. Any failed request that gets retried is billed again in full. If a site returns a challenge page, a rate-limit response, or a partial load, and your scraper retries automatically, you’re paying for the failed attempt and the successful one. A scraper with a high retry rate can double or triple its real bandwidth cost without anyone noticing, because the request count in the dashboard looks the same.
Headless rendering. Running a full browser engine to render a page means downloading everything the page references, because the browser doesn’t know in advance what you actually want. A plain HTTP client that just fetches the HTML and parses it can be a fraction of the transfer size, if the data you need is actually present in the raw HTML rather than injected by client-side JavaScript.
Pagination and infinite scroll. Sites that lazy-load additional content as you scroll will keep firing background requests as long as your automation keeps triggering scroll events. If your scraper doesn’t stop at the exact point it has what it needs, it keeps paying for content it doesn’t use.
Per-GB pricing changes the math
Datacenter proxies are commonly sold as flat-rate or unmetered per IP, so bandwidth is mostly invisible to your budget there, the constraint is usually concurrency and IP count instead. Residential and mobile proxies are usually the opposite: priced per gigabyte transferred, because the provider is paying real device or ISP owners for that bandwidth on the back end.
That structural difference means the same scraping job can have a completely different cost profile depending on which proxy type you route it through. A heavy, image-loaded page routed through a per-GB residential or mobile pool is paying a premium on every byte of page chrome you didn’t need. The same job routed through flat-rate datacenter IPs pays nothing extra for page weight, but may get blocked far more often on sites that specifically filter datacenter ranges, which then costs you in retries instead.
Neither is universally cheaper. The honest answer is that you have to know your own page weight and your own block rate before you can say which pricing model wins for a given target.
Building a real cost model
The model is simple arithmetic once you have real numbers instead of assumptions:
Average transfer per successful page. Pull this from your scraper’s own logs, not from a guess. Measure actual bytes received per completed page for the specific target site you’re scraping, because it varies enormously by site.
Retry rate. What fraction of requests fail and get retried at least once. This is often the number people forget to measure, and it’s the one that quietly inflates bandwidth the most.
Effective cost per page = average transfer per success × (1 + retry rate), multiplied by your price per gigabyte.
Total job cost = effective cost per page × number of pages you need.
Run this before committing to a large job, not after. If the number is higher than expected, the fix usually isn’t a different proxy provider, it’s reducing what each page transfers or reducing how often each page has to be re-fetched.
The levers that actually cut bandwidth
Block unnecessary resource types. If you’re using a headless browser and only need the DOM or specific text content, most browser automation tools let you intercept and abort requests for images, fonts, and stylesheets before they download. This alone often removes the majority of a page’s weight, since visual assets typically dominate transfer size on media-heavy sites.
Prefer the API over the rendered page. Many sites that render data client-side are pulling it from an internal JSON API first. If that endpoint is reachable and its use falls within the site’s terms, fetching it directly instead of rendering the full page can cut transferred bytes by an order of magnitude, because you skip the HTML shell, the JavaScript bundle, and every visual asset entirely.
Cap response size and set real timeouts. A hung connection or a page stuck loading extra assets can sit there accumulating transfer until it times out. Tight timeouts limit the damage from any single bad page.
Fix the retry loop before you scale it. A retry policy that fires immediately and repeatedly on the same failure mode multiplies your bandwidth cost for zero extra data. Back off, cap retry attempts, and log why a request failed before automatically resending it.
Deduplicate before you fetch. Crawlers that re-fetch pages they’ve already visited, or that follow duplicate URLs with different query parameters pointing at the same content, are paying twice for identical bytes.
What blocking does to your bandwidth bill
This is the part that’s easy to miss. Getting blocked isn’t just a failure, it’s an expensive failure. A block usually returns a small challenge page or an error response, so the failed request itself is cheap. The expensive part is what happens next: your scraper retries, gets blocked again, retries again, and every one of those attempts is billed. A high block rate on a site with aggressive bot detection can turn a cheap-looking job into an expensive one purely through retry volume, even though each individual failed request transferred almost nothing.
This is also why “just use better proxies” isn’t a bandwidth fix by itself. No proxy type or rotation strategy is undetectable, and no provider can guarantee a target site won’t flag and challenge your traffic. What actually protects your bandwidth budget is a scraper that behaves predictably, respects reasonable request pacing, backs off instead of hammering, and stops retrying a pattern that’s clearly failing rather than treating every block as a transient error to power through.
The honest bottom line
Bandwidth, not request count, is what a per-GB proxy bill is actually measuring. If you don’t know your average page weight and your retry rate for a given target, you don’t have a real cost estimate, you have a guess dressed up as a number. Measure both before you scale a job, cut the bytes you don’t need with resource blocking and API-first fetching, and treat a rising retry rate as a cost problem worth fixing, not just a reliability annoyance.
If you want more breakdowns like this on proxy selection, pricing structures, and honest tool comparisons, check out the rest of the site here.
Get new guides and videos first — join the Telegram channel.