← all guides

What Your Scraper Leaks Besides Its IP

Most people who get blocked assume it’s the IP. They rotate proxies, the blocks keep happening, and they conclude the proxy provider is bad. Sometimes that’s true. But a lot of the time the IP was never the problem. The request behind it was.

Detection systems built by any site with real traffic worth protecting look at a lot more than where a request came from. They look at how the request is shaped. IP reputation is one signal among many, and it’s often weighted less than people think, because IPs are cheap and rotate constantly on both sides of the fight. What doesn’t rotate as easily is the fingerprint of the client making the request. Here’s what that fingerprint is actually made of.

The TLS handshake talks before you do

Before a single byte of HTTP gets exchanged, a TLS client hello goes out. It lists the cipher suites the client supports, the order it lists them in, the TLS extensions it advertises, and the elliptic curves it’s willing to use. This gets hashed into a fingerprint, commonly known by the JA3 method, and that hash is specific to the TLS library and its configuration, not the website being visited.

A real Chrome browser on a real Windows machine produces one JA3 hash. The requests library in Python produces a different one, because it’s using a different TLS stack (usually OpenSSL through urllib3) with different defaults. Puppeteer with a stealth plugin can spoof the user-agent string in the HTTP layer all it wants, but if the underlying TLS handshake still looks like Node’s default TLS implementation, the mismatch between “claims to be Chrome” and “handshakes like Node” is itself a signal. Detection systems don’t need to catch every scraper this way, they just need to flag the mismatch as one input into a broader score.

This is why swapping proxies never fixes a TLS fingerprint problem. The proxy sits below the TLS layer for a standard HTTP proxy, or it terminates and re-establishes TLS for a transparent one, but either way the fingerprint reflects the client library, not the network path.

HTTP headers carry more information than their values

Header values get spoofed constantly. Header order and formatting get spoofed far less often, because most HTTP clients don’t give you easy control over it. A real browser sends headers in a consistent order specific to its engine, with specific capitalization conventions. Accept-Language capitalized exactly that way, in a specific position relative to Accept-Encoding and Connection.

A lot of scraping libraries send a technically valid but differently ordered header set, sometimes missing headers a real browser always sends (like Sec-Fetch-Site or Sec-Fetch-Mode on modern Chromium), or including headers a real browser would never send in that context. Any of these on their own is weak evidence. Stacked together with a TLS mismatch and an IP with no browsing history behind it, they add up.

HTTP/2 has its own fingerprint

If the connection negotiates HTTP/2, there’s another layer: the SETTINGS frame values, the order streams get prioritized in, and the window update behavior. Real browsers have consistent, engine-specific defaults here too. Most HTTP client libraries used for scraping either don’t support HTTP/2 at all and fall back to HTTP/1.1 (itself a signal, since modern browsers default to HTTP/2 wherever the server supports it), or they implement HTTP/2 with different frame ordering than a browser engine would produce.

The TCP/IP stack underneath all of it

Below TLS, the TCP handshake itself carries a fingerprint: initial window size, TTL, and options ordering in the SYN packet reflect the operating system’s network stack. A residential proxy exit running on a real consumer device inherits that device’s real OS stack. A datacenter proxy or VPS running a headless scraper is usually a Linux box, and its TCP fingerprint reads as Linux server, not Windows or Android consumer device, regardless of what user-agent string the HTTP layer claims. This is one reason residential and mobile exits are harder to distinguish from organic traffic at the network level than datacenter exits are. It’s that the whole stack behind them, not just the IP, resembles a normal device.

Browser-level fingerprints, for anything running headless

If the scraper drives an actual browser engine (Puppeteer, Playwright, Selenium) rather than making raw HTTP requests, there’s a second, much larger fingerprint surface: canvas rendering output, WebGL renderer strings, installed fonts, screen and viewport dimensions, the navigator object’s dozens of properties, timezone versus IP-geolocation consistency, and whether navigator.webdriver and related automation flags are present. Headless Chrome without extra configuration reports itself in ways a plain Chrome install doesn’t. Fixing every individual property is a long list, and missing even one creates a detectable inconsistency, because these properties are cross-checked against each other, not read one at a time.

Timing and cadence

Even with every static fingerprint matched, behavior over time is its own signal. Real users don’t request the same page every 2.0 seconds with no variance. They pause, they click through in an order that reflects reading, and the gap between page load and first interaction has human variance in it. A scraper hitting endpoints on a fixed interval, or fetching pages in an order that skips the navigation a real user would take to get there, produces a session shape that’s statistically distinct from organic traffic even if every individual request looks clean.

What the proxy itself can add

Proxies aren’t just a source of IP addresses, they’re another hop that can leak information if it’s configured carelessly. Some proxy setups add or forward headers like X-Forwarded-For, Via, or Forwarded that expose the original requesting IP or reveal that a proxy is in the path at all. Cheap or misconfigured proxy services sometimes leave these in by default. A proxy that’s actually built for scraping strips them. This is a basic thing to check on any provider you’re evaluating: send a request through the proxy to a header-echo endpoint and read back exactly what arrived. If you see forwarding headers or anything identifying the proxy layer, that’s worth knowing before you rely on it for anything sensitive.

How these get combined

None of these signals on their own proves a request is a bot. Real users run outdated browsers with strange configurations. Real users have proxies too. Detection systems that matter tend to score sessions across all of these dimensions together and flag or challenge sessions that accumulate enough inconsistency, rather than hard-blocking on any single mismatch. That’s also why there’s no single fix. Rotating IPs addresses one input into that score and leaves the rest untouched.

Running a clean stack, practically

For anyone running scrapers in production, the useful takeaway isn’t “how do I fool this system,” it’s “how do I make sure my own stack isn’t accidentally shooting itself in the foot.” A few things worth actually checking, because they’re common causes of avoidable blocks that have nothing to do with proxy quality:

Match your HTTP client’s TLS behavior to what you’re claiming in the user-agent, or use a library that’s built to mirror real browser TLS handshakes if you need HTTP-only requests to look like browser traffic. Keep header sets complete and consistently ordered rather than trimmed down to the bare minimum. If you’re running headless browsers, patch the automation-detectable properties rather than assuming the default headless configuration is invisible, because it isn’t. Build request timing that has some human-like variance in it instead of a fixed interval. And test whatever proxy you’re using against a header-echo endpoint before you trust it in production, since a leaking proxy undermines everything else you’ve done correctly.

None of this makes a scraper undetectable, and no proxy or configuration guarantees a session won’t get blocked or challenged. What it does is remove the easy, avoidable signals, so that whatever detection decision does get made is based on something more meaningful than a sloppy client that gave itself away before the IP even mattered.

If you want the deeper breakdowns on proxy types, rotation setups, and honest comparisons of providers we’ve actually tested, head back to the homepage.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →