← all guides

Why some sites serve different HTML to datacentre IPs (and what it means for a scraper)

You point a scraper at a page, the response comes back with a 200 status, and the HTML is wrong. Prices are missing. The product grid is a skeleton with no data attributes. A block that should hold reviews is empty. Load the same URL from a home broadband connection and it’s all there. Nothing crashed, nothing errored, the site just decided you get a different page. This is one of the more common ways sites treat datacentre traffic, and it’s worth understanding properly before you build anything around it.

What “different HTML” actually means here

There are two distinct patterns that get lumped together as “datacentre IP detection,” and they behave differently.

The first is a hard block: a captcha page, a 403, a redirect to an interstitial. That’s easy to detect and easy to reason about, because your scraper gets a response it clearly isn’t supposed to parse.

The second is quieter and more relevant to this article: the request succeeds, the status code is fine, but the server has decided to serve a reduced or altered version of the page to that IP. Common variants include stripped-out pricing or inventory data, a version of the DOM without the JavaScript hooks that populate content client-side, cached or stale content instead of a live render, or a page that looks complete but has certain fields replaced with placeholder or null values. This second pattern is deliberate cloaking based on IP reputation, and it’s designed specifically to make a scraper think it succeeded when it actually got nothing useful.

How a site knows an IP is a datacentre IP

This isn’t guesswork on the site’s part. There’s a well-established supply chain for this information.

Every IP address belongs to an autonomous system (AS), and that AS is registered to an organisation. AWS, Google Cloud, DigitalOcean, OVH, Hetzner, Linode, and every other hosting provider have their IP ranges publicly listed in regional internet registry records (ARIN, RIPE, APNIC, and so on). A site doesn’t need to do any clever detection work to know a request came from AWS us-east-1; it’s a lookup against a published range. Commercial IP intelligence providers like MaxMind, IPQualityScore, and Spur aggregate these ranges along with reverse DNS patterns (a PTR record ending in .compute.amazonaws.com or .ovh.net is a giveaway) and sell the classification as an API that most anti-bot and CDN products query on every request.

Beyond the base classification, these services also track abuse history for individual IPs and ranges, so a datacentre block isn’t just “this is a server,” it’s often “this specific range has generated flagged traffic before.” That’s why a fresh IP from a hosting provider sometimes behaves differently than one that’s been used for scraping before, even within the same provider.

None of this is specific to any one site’s engineering team building custom detection. Most of it is a checkbox in a CDN or WAF dashboard: Cloudflare, Akamai, and similar platforms let a site owner flag “known hosting/datacentre ASNs” as a category and choose what happens to that traffic, up to and including serving a different backend response.

Why sites do this instead of just blocking

An outright block is loud. It shows up in your logs immediately, and it’s also visible to anyone testing the site manually if they happen to be on the wrong network. Cloaked content is quieter on both ends: your scraper doesn’t throw an obvious error, and the site avoids tipping its hand to whoever is running the scrape about exactly what triggered the response.

The business reasons are usually one or more of: protecting pricing data from competitors who scrape for repricing, protecting inventory levels from being scraped and used to time purchases or resells, reducing load from automated traffic that doesn’t convert, or limiting the value of scraped data to make running a scraper against the site not worth the effort. That last one is the real target of quiet content degradation. A hard block tells a scraper operator “go get different IPs.” A page that looks complete but is subtly wrong can waste far more of someone’s time, because the failure isn’t obvious until the data gets used downstream.

Confirming this is what’s actually happening

Before assuming a site is doing IP-based cloaking, it’s worth ruling out simpler explanations, because they’re more common than people expect. A/B tests, geolocation-based content (currency, language, regional stock), session state, and plain rendering bugs on the site’s end can all produce a difference between two fetches that has nothing to do with your IP’s reputation.

The diagnostic approach here is comparative, not evasive: fetch the same URL from a known clean network and compare the raw response, not just what renders in a browser. Diff the HTML byte for byte where possible. Check whether the difference is in the initial server response or only appears after client-side JavaScript runs, since that tells you whether the cloaking is happening server-side (harder to work around, easier to detect reliably) or is a client-side check reacting to something else. Check response headers too. Some sites set cache-control or vary headers differently for flagged traffic, which is a useful tell independent of the body content.

If the difference correlates specifically with IP classification rather than geography, session, or timing, you’re looking at genuine datacentre-based content variation rather than noise.

What that confirmation should actually change

This is the point where the honest answer isn’t “switch proxy type and keep going.” A site choosing to degrade content for datacentre ranges is making a deliberate policy decision about what it wants automated traffic to see. That’s a signal worth respecting, not just routing around.

The legitimate paths from here are the same ones that apply to any scraping project: check whether the site publishes an API or data feed, which is common for exactly the kind of pricing or catalogue data that gets cloaked (many retailers and marketplaces offer affiliate or partner APIs precisely because they’d rather control the access than have it scraped). Check the robots.txt and terms of service for what the site explicitly permits. If the data is your own listings or something you have a contractual right to pull, request access directly rather than working around the block. And if none of that is available, it’s reasonable to conclude the site doesn’t want to be scraped for this data and to not build a fragile, adversarial pipeline against a site that’s actively trying to stop you.

A note on proxy choice and reputation

Since this is a proxy-focused site, it’s worth being direct about the parts we do talk about: datacentre, residential, and mobile IPs carry different baseline reputations because they come from different kinds of networks, and IP reputation databases treat them differently by default. That’s a fact about how the classification systems work, not a claim about any specific outcome. Residential and mobile ranges belong to ISPs and carriers rather than hosting providers, so they don’t match the ASN lists that flag datacentre traffic. But IP type is one signal among several a site can use. TLS fingerprinting, header ordering, request timing, and behavioural patterns all factor into modern bot mitigation, and none of the proxy categories we cover are undetectable or risk-free against a site that’s investing in detection. Any provider comparison worth reading tells you what was actually tested and what the result was, not a blanket promise.

If you’re evaluating proxy types for a scraping project and want comparisons grounded in what we’ve actually tested rather than marketing claims, that’s what we cover here.

Read more proxy and scraping breakdowns on Proxy Scraping

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →