Why a proxy works in curl but fails in your browser automation
The symptom
You test a proxy with curl. It works. Clean 200, right content, no captcha. You wire the same proxy into Playwright or Puppeteer, point it at the same URL, and get a block page, a captcha wall, or a silent empty response. Same proxy, same IP, same target, different result.
The instinct is to blame the proxy. Usually that is wrong. The proxy moved the bytes correctly in both cases. What changed is everything wrapped around the connection: the TLS handshake, the HTTP headers, the JavaScript environment, and the behavioral signals a real browser produces that curl never does. Sites that care about bot traffic are not just checking your IP address. They are checking whether the thing on the other end looks like a browser all the way down the stack, and curl and a browser look nothing alike to a server, even routed through the identical proxy.
curl and a browser are different clients at the network level
curl is a minimal HTTP client. It opens a TCP connection, does a TLS handshake with whatever defaults its underlying TLS library ships, sends the headers you tell it to send (or a very small default set), and reads the response. That is the entire job.
A real browser does a lot more before a single byte of your page shows up. Chrome, Firefox, and the Chromium engine that Playwright and Puppeteer drive all ship their own TLS stack with a specific, consistent fingerprint: cipher suite order, extensions, supported groups, ALPN settings. This is commonly called a JA3 or JA4 fingerprint. curl’s TLS fingerprint looks nothing like Chrome’s, because it is built on a different TLS library with different defaults.
Here is the part that trips people up: a detection system does not need to see your IP is a proxy to flag you. It can see the TLS handshake shape and immediately know “this claims to be Chrome in the User-Agent header but the handshake says otherwise.” That mismatch alone is a strong signal, independent of what proxy you are on. curl requests usually do not even pretend to be a browser unless you set the User-Agent manually, so on curl the mismatch might not even come up. Point curl’s User-Agent at “Chrome on Windows” and route it through a clean residential IP, and you can still get flagged the moment the TLS handshake does not match a real Chrome handshake, because the two are checked together.
Automated browsers add another layer on top of this. Puppeteer and Playwright, run in default configuration, leave detectable fingerprints in the browser itself: navigator.webdriver set to true, missing or unusual plugin lists, headless-specific rendering quirks, inconsistent screen and viewport values, timing patterns in how JavaScript executes. None of that has anything to do with the proxy. It is the automation framework’s footprint, visible to any page that runs a detection script.
What actually happens when the browser test fails
Walk through what a detection system typically evaluates, roughly in this order, when a request lands:
TLS and HTTP fingerprint. Does the handshake match what the declared browser and OS would actually produce? Does the header order and casing match? curl’s header set and order is not what real Chrome sends by default, and if you have not gone out of your way to replicate a browser’s header ordering in your automation, Playwright’s default headers may also not perfectly match a stock Chrome install depending on version and configuration.
IP reputation. This is the layer most people think is the whole story. Is the IP a known datacenter range, a known proxy pool, or does it look like a residential or mobile subscriber? This is real and it matters, but it is one signal among several, not the only one.
JavaScript environment checks. Once a page loads (or before it fully renders, via a script injected early), sites can probe navigator properties, canvas and WebGL fingerprints, font lists, and automation-specific tells like navigator.webdriver. curl never executes JavaScript, so this entire category of check simply does not apply to a curl request. It is the single biggest reason a curl test and a browser test are not comparable: curl is invisible to an entire layer of detection that browser automation walks straight into.
Behavior. Mouse movement, scroll patterns, timing between actions, whether a captcha challenge gets solved. curl produces none of this because it does not render anything. A browser produces all of it, and if the automation moves in obviously scripted patterns (constant velocity mouse paths, zero-latency clicks, identical timing between every action), that is itself a signal.
A proxy provider selling “block-free scraping” is selling you the IP layer only. It cannot fix your TLS fingerprint, it cannot make Puppeteer stop reporting navigator.webdriver, and it cannot make your click timing look human. If the browser test still fails after confirming the proxy itself is clean, the proxy is very likely not the layer that is broken.
How to actually isolate the cause
Do not guess. Test each layer separately before concluding it is the proxy’s fault.
-
Confirm the proxy’s IP is clean on its own terms. Hit an IP reputation lookup or your own test endpoint through the proxy with curl. Check whether it is flagged as datacenter, whether it shows up on known proxy or VPN lists, and whether the ASN matches what the provider claims to be selling you (residential, mobile, or datacenter).
-
Check what your browser automation is actually sending. Route Playwright or Puppeteer through the proxy to a header-echo endpoint (something as simple as httpbin’s headers endpoint) and compare the full header set, order, and User-Agent against what a manually operated copy of the same browser sends to the same endpoint. Differences here point at the automation layer, not the proxy.
-
Check the TLS fingerprint separately. There are public JA3/JA4 fingerprint test endpoints. Hit one with curl through the proxy, then hit the same one with your automated browser through the proxy, and compare. If curl’s fingerprint and the browser’s fingerprint land in different buckets from what the target site expects for a “real Chrome,” that mismatch exists regardless of proxy quality.
-
Check for automation tells in the JS environment. Load a fingerprinting test page in your automated browser and look specifically at
navigator.webdriver, plugin count, and headless-specific quirks. This is entirely on the automation configuration, not the network path. -
Only after 1 through 4 come back clean, treat the proxy as a suspect. If the IP itself checks out, the headers match, the TLS fingerprint matches, and the JS environment looks like a normal browser, and the block still happens only with that specific proxy, then something about the proxy’s routing, latency, or IP history specific to that target site is worth investigating with the provider.
What this means for choosing a proxy
None of this means IP choice is irrelevant. It means IP choice is necessary but not sufficient. A residential or mobile IP from a reputable pool solves the IP reputation layer. It does nothing for TLS fingerprinting, browser automation fingerprints, or behavioral signals, because those live entirely in your client stack, not the proxy.
If your workflow requires a real rendered browser (JavaScript-heavy pages, content behind client-side rendering, interactions that require an actual DOM), you are going to be evaluated on all four layers above, and the proxy only ever answers one of them. If curl is sufficient for your target (a plain HTML response, an API endpoint, static content), you skip the TLS-fingerprint-mismatch and JS-environment problems entirely just by not needing a rendered browser, which is one reason curl tests look deceptively easy to pass compared to the same target through automation.
No proxy, technique, or fingerprint-spoofing library makes a request undetectable or guarantees a scraper stays unblocked. Detection systems evolve, and what passes today is not a permanent state. Any scraping should stay inside the target site’s terms of service and applicable law, and avoid personal or paywalled data. The point of understanding these layers is not to defeat detection, it is to know honestly which part of your stack is actually causing the block you are seeing, instead of blaming the proxy for a problem that lives in your TLS library or your automation framework’s default fingerprint.
If you want to go deeper on how residential, mobile, and datacenter proxies actually differ, and where each one fits a scraping workflow, browse the rest of the guides on proxyscraping.org.
Get new guides and videos first — join the Telegram channel.