← all guides

What TLS Fingerprinting Reveals About Your Scraper

The handshake happens before your proxy matters

Every scraping conversation eventually turns into a conversation about IPs: residential versus datacenter, rotation intervals, ASN diversity. That conversation matters, but it starts one layer too late. Before your request ever reaches the site’s application logic, before a single cookie is set, your client already had a conversation with the server at the TLS layer, and that conversation left a fingerprint that has nothing to do with your IP address.

When your scraper opens an HTTPS connection, it sends a ClientHello message. That message lists the TLS version it supports, the cipher suites it offers, the extensions it understands, the elliptic curves it can use, and the order it puts all of that in. Real browsers have a very specific, very consistent way of building this message, because it is generated by the browser’s own TLS stack (BoringSSL for Chrome, NSS for Firefox). Most HTTP client libraries used in scraping, requests, urllib3, aiohttp, Go’s net/http, use a different TLS stack entirely, and they build the ClientHello differently. That difference is fingerprintable, and it survives a proxy swap completely intact.

JA3 and JA4, in plain terms

JA3 is the fingerprinting method most detection vendors reference. It takes the ClientHello fields, TLS version, cipher list, extension list, elliptic curve list, elliptic curve point formats, concatenates them in order, and MD5-hashes the result. The output is a 32-character hash that represents “how this specific TLS implementation, in this specific configuration, says hello.” Two connections from two different IPs, using the same underlying HTTP client with default settings, produce the same JA3 hash every time.

JA4 came later as an attempt to fix some of JA3’s weaknesses, mainly that JA3 orders extensions the way the client sent them, which browsers with extension-shuffling (a defense some browsers added) can scramble intentionally. JA4 normalizes some of that ordering and adds more structure to the hash format, so it is a bit more resistant to trivial reordering tricks. Both methods answer the same underlying question for a site operator: does this handshake look like a real browser’s TLS stack, or does it look like a library’s?

The reason this matters operationally is that JA3/JA4 hashes for common scraping libraries are widely published and widely known. requests with default urllib3, Python’s httpx, Go’s default HTTP client, Scrapy’s default downloader, all of these have well-documented, static JA3 hashes. A site does not need to see your traffic pattern or your IP reputation to flag you. It can flag the handshake alone, on the very first packet, before your request even carries a User-Agent header.

Why IP rotation does not fix this

This is the part that trips up a lot of scraping setups that otherwise look solid. You buy residential proxies, you rotate every request, you diversify ASNs, you match geolocation to the target audience, and you still get blocked or served degraded content. The IP layer is clean. The TLS layer is not.

Rotating your proxy changes where the connection appears to originate. It does nothing to change how the connection is negotiated. If your HTTP client is still building the same ClientHello every time, a detection system that fingerprints at the TLS layer sees the same signature walk in from a hundred different IPs in a row. That pattern, one static handshake, many rotating source addresses, is itself a signal, and it is often a stronger signal than a single suspicious IP would be. Some detection stacks weight it that way explicitly: a JA3 hash associated with automation tooling, seen across a wide spread of IPs in a short window, reads as a distributed scraping operation rather than a diverse population of real users.

This is also why “just use a good proxy provider” is incomplete advice. A residential proxy fixes IP reputation and ASN classification. It does not touch anything upstream of the socket. The TLS handshake is generated by your client’s networking stack, not by the proxy, and a proxy that tunnels your traffic (as opposed to terminating and reissuing the TLS connection itself) passes your original ClientHello straight through unchanged.

What actually changes a TLS fingerprint

A handful of approaches exist for matching a scraper’s TLS signature to a real browser’s, and it is worth being precise about what each one actually does, because they are not interchangeable.

Browser automation (Playwright, Puppeteer, Selenium) drives a real browser binary, so the TLS stack doing the handshake is the actual browser’s TLS stack. This is the most faithful match by construction, because there is no separate library trying to imitate anything. The tradeoff is resource cost: a real browser process is heavier than an HTTP client, and running hundreds of concurrent browser instances is a different infrastructure problem than running hundreds of concurrent HTTP requests.

TLS impersonation libraries (curl-impersonate, and various language ports of the same idea) reimplement a specific browser’s ClientHello construction inside a lightweight HTTP client, without running the browser itself. These can produce a JA3/JA4 hash that matches a real Chrome or Firefox release closely, sometimes exactly, for that specific version. The catch is version drift: browsers update their TLS stacks, and an impersonation library pinned to an old Chrome fingerprint will start looking anomalous again once that Chrome version ages out of the real-world population, or once a site’s detection starts checking the fingerprint against other signals, like the TLS version claimed versus the HTTP/2 settings frame that follows it, which impersonation layers do not always keep in sync.

Default HTTP clients with no adjustment present their library’s native fingerprint, unmodified, on every single request. This is the baseline most scraping setups run without realizing it, and it is the easiest case for a detection system to classify.

None of these options make a scraper undetectable. They change what one specific signal reveals. A site with a mature detection stack correlates TLS fingerprint with HTTP/2 header ordering, JavaScript execution signals, request timing, and behavioral patterns, so matching the TLS layer removes one input to that model, not the model itself.

What this means for how you build

If you operate scrapers at any real scale, treat TLS fingerprint as a first-class configuration decision, not an afterthought you reach for after getting blocked. Know what your HTTP client’s default fingerprint looks like before you deploy it. If you are running browser automation, keep the browser binary and its automation driver on a current, matched version pair, since mismatched versions between the driver and the browser can themselves produce an inconsistent handshake. If you are running a lightweight HTTP client for volume, decide deliberately whether TLS-layer detection is a factor for your specific targets, because plenty of sites do not check it at all, and paying the overhead to defeat a check that was never running is wasted engineering.

And keep the proxy and TLS layers mentally separate. A good proxy gets your IP reputation and geolocation right. It says nothing about your handshake. Judge each layer on what it actually controls, and be honest with yourself about which layer is causing a given block before you spend money fixing the wrong one.

None of this is a guarantee against blocking, and no configuration here makes a scraper invisible. Sites that care about automated traffic combine TLS fingerprinting with several other signals, and a clean handshake just removes one easy tell, not the whole detection surface. Always scrape within the target site’s terms of service and applicable law, and treat any technical adjustment as a matter of building a technically honest client, not a workaround for restrictions the site has put in place deliberately.

If you want more breakdowns like this on how scraping infrastructure actually works under the hood, head back to the Proxy Scraping home page for the rest of our guides on proxies, detection, and scraping tools.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →