Budgeting bandwidth for image heavy scrapes
Most scraper budgets get built around request counts. Pages per day, requests per minute, proxies per pool. That works fine for text-heavy targets like listings pages or search results. It falls apart fast on anything with a product gallery, a photo feed, or a map tile layer, because the thing that actually costs money on those targets isn’t the request, it’s the bytes.
If you’ve run a scrape against an e-commerce catalog or a real estate site and watched your proxy bill triple overnight, this is usually why. Here’s how to think about bandwidth as its own budget line, separate from request volume, and what actually moves the number.
Why proxies bill in gigabytes, not requests
Datacenter proxies are usually sold as a fixed number of IPs, sometimes with unlimited or high-cap bandwidth, because the provider’s cost is the server and the IP lease, not the traffic itself. Residential and mobile proxies work differently. The provider is paying real people or a mobile carrier for the exit bandwidth their device or SIM pushes, so the pricing model follows that cost structure directly: you buy gigabytes, and every byte that goes through the proxy, in both directions, comes out of that pool.
This is the part that catches people out. A scraper built and tested on a datacenter pool, where a 2MB product image costs nothing extra, behaves completely differently once it’s pointed at a residential pool that’s metered per GB. The code didn’t change. The bill did.
Where the bytes actually go
Text-only pages are small. A rendered HTML document for a product page is typically tens of kilobytes. The images on that same page are usually the majority of the page weight by a wide margin, sometimes ten to fifty times heavier than the markup itself, because a single product gallery page can load a primary image, several thumbnails, a zoom variant, and a handful of related-product thumbnails, all before you’ve extracted a single price.
A few specific sources of bandwidth waste show up over and over in image-heavy scrapes:
- srcset and responsive images. Modern sites serve multiple resolutions of the same image for different screen sizes. A headless browser rendering the page like a real device will often pull more than one of these variants.
- Lazy-loaded images below the fold. If your scraper scrolls the page or waits for network idle before extracting data, it can trigger the load of images you never needed.
- CDN retries on flaky proxy exits. A slow or unstable residential connection that times out mid-download and retries doesn’t just waste time, it re-downloads the asset from zero, effectively doubling the cost of that one image.
- Uncompressed or high-quality image requests. Some sites serve a different, heavier image to clients that don’t send the right
Acceptheader or don’t support modern formats like WebP or AVIF, which a poorly configured scraper can trigger without meaning to.
None of this is exotic. It’s just how normal web pages are built, and a scraper that behaves like a full browser inherits all of it.
The retry multiplier
Rotating proxies make bandwidth budgeting harder in a specific way: every failed request costs bandwidth too, and then gets paid for again on the retry. A request that gets a partial response, a timeout, or a block page still transferred data before it failed. If your retry logic just fires the same request again through a new IP without checking what actually failed, a single flaky page can cost two or three times its real weight before you get usable data back.
This is one of the honest arguments for keeping proxy pools smaller and higher quality rather than maximizing IP count. A pool with a high failure rate doesn’t just slow a scrape down, it inflates the bandwidth budget through repeated partial transfers that never resulted in usable output.
The concrete levers that actually save bandwidth
Turn off image rendering in headless browsers. If you’re using Playwright or Puppeteer and you only need the DOM or the text content, you don’t need the browser to actually fetch images at all. Both support request interception, where you can abort any request whose resource type is image before it leaves the browser. This is the single biggest lever available for image-heavy targets, because it removes the image bytes entirely instead of trying to shrink them.
Use the network layer instead of a browser when you can. If a target doesn’t require JavaScript execution to expose the data you need, a plain HTTP client like requests or Scrapy’s downloader never renders a page or fetches embedded assets in the first place. Scrapy in particular only pulls the URLs you explicitly request, so image bandwidth only shows up if your parsing code follows those URLs.
Pull thumbnail or CDN-resized endpoints when the target supports them. Many sites serve images through a CDN that accepts size parameters in the URL, for a listing thumbnail versus a full-resolution photo. If you only need to confirm an image exists, or extract EXIF-level metadata rather than the visual content, requesting the small variant instead of the original can cut that asset’s weight by an order of magnitude.
Use HTTP range requests for partial reads. If you genuinely need something like image dimensions or file headers rather than the full image, a range request (Range: bytes=0-1023) can pull just the header bytes of most image formats instead of downloading the whole file. This only works if you’re reading the file structure directly rather than relying on a browser or an image library that expects the full file.
Cache and deduplicate. Product photos and listing images often don’t change between scrape runs. Hashing the image URL (or a stable identifier from the page) and skipping re-download on unchanged assets avoids paying for the same bytes twice. This matters more the more frequently you re-scrape the same target.
Respect Content-Length before committing to a download. For clients that support it, checking the response header before pulling the full body lets you skip unexpectedly large files rather than discovering the size after the transfer is already metered.
Building the actual budget number
Once you know where the bytes go, the budget itself is arithmetic. Take an honest estimate of pages per run, multiply by an honest estimate of average page weight including images (measured, not guessed, ideally from a small pilot run against the real target), then apply a retry multiplier based on the expected failure rate of the proxy pool you’re using. Compare that number against the provider’s pricing tiers in gigabytes, not in IP count or request count, since that’s the unit that will actually determine the bill on a residential or mobile pool.
It’s worth running that pilot on a small batch of real pages before committing to a full run. Page weight varies a lot by site, and a guess based on a “typical” page is often off by several times in either direction once real images, retries, and lazy loading are accounted for.
Monitoring while it runs
A budget built before the run doesn’t help much if nothing is watching bandwidth while the scrape is live. Most proxy providers expose per-session or per-proxy usage stats through their dashboard or API. Logging bytes transferred per request, alongside the URL and response status, makes it possible to catch a single misbehaving session, one that’s stuck in a retry loop or accidentally pulling full-resolution images, before it consumes the whole month’s allocation. This is a cheap check to build in and it’s the difference between finding a runaway session in an hour versus finding it at the end of a billing cycle.
Bandwidth-heavy targets aren’t a reason to avoid residential or mobile proxies when the target genuinely needs them for legitimate access. They’re a reason to measure page weight before scaling, decide up front what you actually need out of each page, and strip the rest before it ever leaves the proxy.
For more on choosing proxy types for specific scraping workloads and reading providers’ bandwidth pricing honestly, head back to the proxyscraping.org homepage.
Get new guides and videos first — join the Telegram channel.