← all guides

Scraping a site that builds itself in the browser

web-scraping headless-browsers javascript-rendering bandwidth proxies

I inherited a scraper that moved about 400 GB a month. After the rewrite it moved 9.

Same source, same pages, same rows landing in the same table. The old version opened a real browser for every URL. The new one asks the site for the data that browser was going to fetch anyway.

I sell residential and mobile proxy lines, and residential is billed by the gigabyte, so every word of this costs me money. Weigh it accordingly.

Three costs, and the one that actually hurts

A plain HTTP request holds a socket, a buffer, and whatever your parser needs for a moment. A few hundred KB of memory. One core will drive hundreds of them at once.

Chromium with a single page open sits at 300 to 500 MB before that page has done anything interesting. Add layout and script execution and a box that would happily carry 400 plain connections in flight will manage six or eight browser instances.

Memory and CPU are the costs people see, because they can hear the fan.

Transfer is the one that turns up on an invoice.

What a rendered page actually pulls down

You asked for one document. You get that document plus everything it references: images, webfonts, a couple of stylesheets, an analytics script, a tag manager, a chat widget, a consent banner, and however many ad vendors that page loads this week.

Count the requests on an ordinary retail product page. 80 to 140 is normal. One of them is yours.

I measured one target properly last year because the bill had stopped making sense. Rendered in full, a product page cost 2.6 MB across the wire. The JSON behind that same page was 31 KB.

Eighty to one, on a line billed per gigabyte.

Push 60,000 pages a night through that and the rendered version moves 156 GB. The direct version moves under 2. Nobody argues about the second invoice.

It is why people scraping from unmetered datacenter IPs get such a shock when a target forces them onto residential. Same job, same code, and the transfer they never had to think about becomes the biggest line item they have.

Open the network tab before you open a browser

Load the page in your own browser with devtools on the network tab. Reload. Filter to Fetch/XHR, sort by size, and read the response previews from the top down.

A page that assembles itself in the browser has to get its content from somewhere. The javascript did not invent the price. It arrived over the network, structured, from an address sitting right there in the list.

That source is smaller than the rendered page, and better labelled.

Smaller because it is data with none of the furniture. Better labelled because the fields already have names, so you stop maintaining CSS selectors that break the week somebody renames a class.

It also moves less. That endpoint is what the site’s own front end runs on. Redesigns shuffle markup constantly. Rewriting the data layer means rewriting their app, so it happens far less often.

When you find one, use Copy as cURL in devtools. That hands you the working request with every header and cookie attached, and you can then strip it back one header at a time until it breaks. What is left is the minimum request that works.

Sometimes it is already in the HTML

A lot of frameworks serialise the whole initial page state into a script tag and ship it with the document. Next.js parks it under an id you can search for. Others hang it off a global.

View source, search for a value you can see rendered on the page, and if it turns up inside a data structure you are finished before you started. One request, no browser, already parsed.

That check takes about thirty seconds, so it goes right after the network tab.

When a browser is genuinely the answer

Two cases where none of the above helps.

Content that only exists after an interaction. A filter that has to be set, a tab that has to be clicked, a scroll that triggers the next batch, where the fetch behind it carries a token the page mints at runtime.

And a challenge that has to execute. A site putting a javascript challenge in front of everything is telling you what it wants, and going around that is a question about its terms rather than a technical puzzle. Read the terms and the robots file first. If the answer there is no, it is no.

There is a third case that is about your calendar rather than the site. Deadline Friday, endpoint behind a signing scheme, browser works today. Fine. Log the decision somewhere, because the cost of it recurs monthly and the reason for it does not.

Running a browser thin

Block by resource type. Every automation library can intercept and abort requests. Kill images, media, fonts and stylesheets and the DOM still holds every word of text you came for. On a heavy page that is 70 to 80% off the transfer, and it takes work off the CPU too, because nothing has to decode a picture nobody will look at.

Block third party hosts as well. Analytics, tag managers, ad calls, session recorders. None of it touches your data and all of it lands on your bill. I keep a hostname list for this that barely changes between targets.

Then stop starting a process per page. Plenty of scrapers launch a browser, load one URL, close it, and repeat that 100,000 times. Startup is one to two seconds of pure overhead each time, which is over a day of wall clock across that job.

Run one instance and open a fresh context per unit of work. Contexts carry their own cookies and storage, so you keep the isolation and skip the cold start.

Do not overload the instance in the other direction either. Forty tabs on one browser runs out of memory in a way that fails slowly and lies to you about the reason. A few pages per instance, a few instances, restarted on a schedule, because they leak.

Waiting is where the flakiness lives

Most of the failures people blame on the target come from how they wait.

Waiting for the network to go idle is the default and the worst option available. Any page with a poller or an open websocket never goes idle, so you eat the full timeout on every single page. Multiply 30 seconds by 100,000 and see what you get.

Fixed sleeps fail the other way. Two seconds is wasteful on a fast day and too short on the slow day, which is the day you needed it to hold.

Wait for the specific element, or for the specific response landing on the network. That returns the moment the thing exists, usually a few hundred ms instead of five seconds of watching an idle timer.

Five seconds a page against 800 ms, over 100,000 pages, is the gap between finishing overnight and not finishing.

The failures get honest as well. If the selector never resolves, that is a real event worth logging, and it usually means the page said something you were not reading. An empty result set. A notice. A form.

A heavier client shows more of itself

There is a belief that a real browser is the safer choice because it looks more like a person.

Half true. A browser produces a browser’s headers and a browser’s TLS handshake, and that removes an entire category of mismatch between what a client claims and what it sends.

The other half is that an automated browser puts far more surface on the table. A plain client sending eight headers gives a site eight things to look at. A headless browser under automation reports a rendering engine, a font list, a screen size, a timezone, a pile of hardware answers, and the timing of everything it fires, and every one of those can be checked against the others.

I am not going to tell you how to file any of that down. The narrow point is that paying 10x per page does not buy you safety, so do not pick the expensive option for a benefit it was never going to deliver.

The one I built the expensive way

A marketplace source, built as a browser fleet. Eight instances across two boxes, a queue in front, restart logic, health checks. Two weeks to get stable, then five months in production at roughly 190 GB a month on residential addresses, on a machine that did nothing else.

One slow afternoon I had that site open in my own browser with the network tab up for a completely unrelated reason. One request. Every field I had been prising out of the rendered page, already named. No token, no signature, paging as a query parameter.

The rewrite took an afternoon. Transfer dropped to about 4 GB a month and the box went back to doing other work.

Five months, and I never once opened the network tab, because I decided on day one that it was a javascript site and that decision closed the question before I asked it.

Devtools now opens before I write a line, on every new source.

What this does not tell you

Whether your target has a usable source behind it. Plenty do not. Some sign every request, some bind it to a session the page mints, some genuinely assemble the content client side out of fragments that only mean anything together.

And finding one is not permission to hammer it. Terms and robots still apply, rate limits still apply, and an endpoint being easy to call does not mean you get to call it 50 times a second.

The block list I load on every job, and the rest of what I run on my own lines, is here.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →