Cloudflare Turnstile in scrapers: what actually passes
if you run scrapers for a living, you have hit the little Turnstile checkbox, or worse, the version you never see that quietly decides your request is not worth serving. it does not look like the old CAPTCHAs. there is no traffic light grid to farm out. the widget runs a bundle of browser checks, scores the session, and hands back a token or it does not. when it does not, your pipeline stalls on a 403 or a page that never finishes loading, and the error log tells you almost nothing.
the stakes are mostly about cost and predictability. a scraper that passes 95% of the time on a clean test run and 40% of the time at volume is a scraper that eats proxy bandwidth, retries, and engineer hours. I am based in Singapore and run a mix of mobile, residential and datacenter traffic for my own projects, and most of what follows comes from watching what survives contact with real targets, not from vendor marketing.
this piece is for people who already know what a headless browser is, what a residential proxy is, and why a plain requests call gets blocked. I will cover what Turnstile actually checks, what passes and what does not, three worked examples with arithmetic you can adapt, the failure modes that bit me, and a short list of operator notes. nothing here is legal advice, and I will say more on that near the end. if you want the wider context first, the blog index has the proxy basics.
background and prior art
Cloudflare launched Turnstile publicly in September 2022 as a free alternative to CAPTCHA, positioned as privacy-preserving and low friction. the announcement on the Cloudflare blog describes the idea: run a rotating set of non-interactive browser challenges, and only fall back to a human-facing step when the signals are ambiguous. site owners embed a JavaScript widget, the widget produces a token, and the site’s backend verifies that token against Cloudflare before accepting a form post, login or API call. it is separate from, and often deployed alongside, Cloudflare’s bot management and the “Just a moment” interstitial, which people constantly confuse with it.
the prior art on the scraping side is a decade of CAPTCHA solving: image labelling farms, then reCAPTCHA v2 audio tricks, then token-farming services that proxy a challenge to a human or a model and return the answer. Turnstile changed the economics because there is often nothing to solve visually. the check is about whether your environment looks like a real browser on a real network. there is also a standards angle. Private Access Tokens, described in RFC 9576, let a device attest itself to an issuer without revealing identity, and Cloudflare has supported them for Apple devices. that matters because the direction of travel is clear: the cheap signal is “this is genuine hardware attested by its vendor,” and that is something a scraper farm cannot fake with a patched Chromium.
the core mechanism
it helps to separate three things that get called “Turnstile.”
- the widget: a script loaded from
challenges.cloudflare.comthat renders in an iframe, runs checks, and writes a token into a hidden form field namedcf-turnstile-response. - the token: an opaque string issued after the widget is satisfied. per the Turnstile docs, tokens are single use and expire after 300 seconds.
- the verification: the site’s backend posts the token to
https://challenges.cloudflare.com/turnstile/v0/siteverifywith its secret key, and Cloudflare answers success or failure. the server-side validation docs cover the response fields, including the hostname the token was issued for and error codes such astimeout-or-duplicate.
the widget has three modes, set by the site owner: managed (Cloudflare decides whether to show a checkbox), non-interactive (a spinner, no click), and invisible (nothing visible at all). from a scraper’s point of view the mode matters less than you would think. in all three, the real work is the same background evaluation, and the visible click is the cheap part.
what does the evaluation look at? Cloudflare does not publish the full list, and anyone who tells you they know the exact weights is guessing. from the public docs and from observed behaviour, the inputs fall into four buckets:
- network: the IP’s reputation and type, the ASN, the TLS fingerprint (JA3/JA4 style), HTTP/2 settings frames and header order. a Chrome-looking user agent arriving over a Python TLS stack is a mismatch that shows up before any JavaScript runs.
- browser environment: the JS engine’s behaviour,
navigatorproperties, WebGL and canvas output, fonts, timing of certain API calls, and whether automation artefacts exist (navigator.webdriver, CDP runtime side effects, patched functiontoStringoutput). - behaviour: pointer and keyboard events, timing between page load and token request, whether the page was interacted with at all.
- session consistency: does the claimed OS, screen size, language, timezone and IP geography tell one coherent story. a US-English Windows profile with a Jakarta residential IP and a UTC timezone is not impossible, but it is a pattern.
the token is the output, and this is the point that trips up most first attempts: a token is bound to the sitekey and the hostname it was issued for, it is single use, and it dies in five minutes. you cannot harvest a pile of tokens once and replay them. each protected action needs a fresh token, issued in a context the site’s backend finds plausible.
a minimal client side render looks like this, and it is worth reading the client-side rendering docs if you have not:
<script src="https://challenges.cloudflare.com/turnstile/v0/api.js" async defer></script>
<form action="/login" method="POST">
<div class="cf-turnstile" data-sitekey="YOUR_SITEKEY"></div>
<button type="submit">Sign in</button>
</form>
and the verification on the site’s backend, which is what your solved token eventually has to survive:
curl https://challenges.cloudflare.com/turnstile/v0/siteverify \
--data "secret=SECRET_KEY" \
--data "response=TOKEN_FROM_FORM"
two practical consequences. first, if you do not control the target, you never see the verification step, so a “valid looking” token that the backend rejects shows up only as the login or form failing. second, Cloudflare provides dummy sitekeys for testing (for example 1x00000000000000000000AA always passes), which is useful for checking your own harness without hitting a live target. use those to test plumbing, then stop. they say nothing about whether you pass real evaluation.
what passes, in order of reliability
here is the ranking as I would give it to a colleague, from the approaches I have run myself.
- a real, unpatched, headful browser on a clean residential or mobile IP, with human-plausible timing. this passes most often. it is also the slowest and most expensive per page.
- a browser driven by an automation framework built to hide CDP and webdriver artefacts, on a good IP. this passes often but is a moving target. what works this month can fail after a Cloudflare update.
- a solver service that returns tokens from its own browser farm. this passes when the farm’s environment is good, but you have little visibility and the token has to be used from a context the backend accepts.
- a stock headless Chromium or Playwright with default settings. this fails often enough that I do not budget for it on protected targets.
- raw HTTP clients with spoofed headers. this fails on any page where Turnstile gates the action, because there is no JS execution to produce a token at all.
none of that is a guarantee, and I would distrust any benchmark that claims otherwise, including mine.
worked examples
the numbers below are arithmetic on stated assumptions, not measurements from a public benchmark. swap in your own pass rates and prices. I am deliberately not quoting solver prices, because they change and vary by provider; check the current price list of whichever service you evaluate.
example 1: a form-gated listing scrape, 100,000 pages
setup: a directory site shows listing pages freely, but every 10th page request triggers a Turnstile check before it serves more. you need 100,000 listing pages.
- pages per challenge: 10, so roughly 10,000 challenges for the run.
- approach A, real browser with a persistent profile: once a session passes, Cloudflare issues a clearance cookie scoped to the site, and subsequent requests in that session ride on it. so instead of 10,000 challenges you may see far fewer, because a warm session may not be challenged again for a while. the cost is browser memory and a sticky IP.
- approach B, token solver per challenge: 10,000 solves. if your provider charges a flat price per thousand, the bill is 10 times that price. at an assumed 90% first-try pass rate you pay for about 11,111 attempts if failed ones are billed, fewer if they are not. read the billing terms for failed tasks.
the lesson from this one: how many challenges you face depends heavily on whether you preserve session state. a scraper that throws away cookies between pages multiplies its own challenge count. I have seen the same job go from thousands of challenges to a few hundred just by reusing a warm context per sticky IP for the cookie lifetime.
example 2: login flow with a per-attempt token
setup: you need to log in to an account you own on a portal protected by Turnstile on the login form, then fetch 500 pages of your own data.
- the token is needed once per login, not per page. 500 pages after one login means one token.
- sequence: open the page in a real browser context, let the widget run, wait for the hidden field to populate, submit. total wall time is typically a few seconds when it passes.
- if you try to be clever and pre-solve ten tokens at the start of the day, you will find them dead by the time you use the last one. the 300 second expiry and single-use rule make batching pointless. generate the token immediately before the action that consumes it.
# playwright, headful, one context per account, sticky proxy
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=False)
ctx = browser.new_context(
proxy={"server": "http://PROXY_HOST:PORT"},
locale="en-SG",
timezone_id="Asia/Singapore",
)
page = ctx.new_page()
page.goto("https://portal.example/login")
page.wait_for_selector('input[name="cf-turnstile-response"]', state="attached")
page.wait_for_function(
"document.querySelector('input[name=\"cf-turnstile-response\"]').value.length > 0",
timeout=30000,
)
page.fill("#email", "[email protected]")
page.fill("#password", "...")
page.click("button[type=submit]")
note that locale and timezone are set to match the proxy’s geography. that single change fixed more of my intermittent failures than any stealth plugin did.
example 3: mobile versus datacenter IP, same browser
setup: identical headful browser profile, identical script, 200 login-page loads against a Turnstile-protected form, run once through a datacenter range and once through a mobile carrier IP.
I will not quote a pass rate I did not record in a form I can publish. what I can say reliably is the direction: for the same environment, the IP class changed the outcome more than any browser tweak, and mobile carrier IPs, which sit behind carrier-grade NAT and are shared by many real users, are the hardest for a site to blanket-penalise. the cost is bandwidth price. if your pages are 300 KB each, 200 loads is about 60 MB, which on a per-GB mobile plan is cheap for a test and meaningful at 10 million pages. do the multiplication before you commit. the earlier post on picking a target country for a scrape covers why geography of the exit IP interacts with these checks.
edge cases and failure modes
these are the five that cost me the most time.
1. the token passes, the action still fails
the token is valid but the site rejects the form. causes I have seen: the token was issued for a different hostname than the one posting it (a subdomain mismatch), the token was reused (timeout-or-duplicate), or the backend also checks something else, such as a session cookie that was never set because you skipped a page. counter-strategy: replay the exact human flow in a browser first, record what cookies and requests precede the post, and only then try to shorten it.
2. warm session, cold IP
you keep a good browser profile with a clearance cookie, then your proxy rotates the exit IP mid-session. the cookie is tied to more than the browser, and the IP change invalidates trust. counter-strategy: use sticky sessions and rotate per account or per job, never mid-flow. if your provider’s sticky window is shorter than your flow, fix the flow, not the cookie.
3. stealth patches that make you more unusual
a pile of JavaScript patches meant to hide automation can leave their own fingerprints: functions whose toString output differs from native, property descriptors in the wrong order, an overridden navigator object that fails identity checks. I have seen scrapers get worse after adding a stealth plugin. counter-strategy: start from the least modified setup that works, add one change at a time, and keep a control profile with no patches to compare against.
4. concurrency makes you look like a farm
one browser per IP behaves well. fifty browsers sharing one IP, all requesting tokens within the same second, do not. timing correlation across sessions is cheap for Cloudflare to see. counter-strategy: cap concurrency per IP at one or two, add jitter to start times, and spread across subnets and ASNs when you can.
5. solver services as a hidden dependency
a solver returns a token generated elsewhere. the environment that produced it (IP, fingerprint, timezone) differs from the one that submits it, and some sites tie the two together. the other risk is operational: the provider’s success rate drifts when Cloudflare changes something, and you find out from your own failures. counter-strategy: keep a second path (your own browser) behind the solver, log pass rate per hour, and alert on a drop instead of trusting a monthly average.
a general rule across all five: log the outcome of every challenge with the IP, the profile and the time. without that data, every failure is an anecdote.
what we learned in production
the largest single lever was not stealth, it was reducing how many challenges I trigger. preserving cookies, reusing warm contexts on sticky IPs, spacing requests like a person who reads the page, and not retrying instantly on a failure all cut challenge volume. each avoided challenge is a challenge I do not have to pass, and that beats any bypass technique on both cost and reliability. the second lever was matching the whole story: IP geography, locale, timezone, language header and screen size all pointing at the same place.
the other lesson is to treat this as maintenance, not a project you finish. Turnstile is updated continuously, and a setup that passes today may degrade without any change on your side. I keep a small canary job that loads a Turnstile-protected page I am entitled to access every hour and charts the result, so I notice drift before a client does. if you work across several account-heavy workflows, the same discipline applies in adjacent fields: the sister site multiaccountops.com covers session isolation and profile hygiene, and antidetectreview.org reviews the browser tools that people use for that. for the proxy side, the blog index has the type-by-type breakdown.
one last point on boundaries. only scrape what you are allowed to access, respect the target’s terms and robots rules where they apply, and do not use any of this to bypass access controls on data you have no right to. the legal position on scraping and on circumventing technical measures varies by country and by case. this is not legal advice, and if the stakes are real, talk to a lawyer in the jurisdiction that matters.
references and further reading
- Cloudflare Turnstile documentation: widget modes, token lifetime, test sitekeys.
- Turnstile server-side validation: the siteverify endpoint, response fields and error codes.
- Cloudflare blog, Turnstile announcement: the stated design goals and privacy position.
- RFC 9576, The Privacy Pass Architecture: the standards background for device attestation tokens such as Private Access Tokens.
- Turnstile client-side rendering: how the widget is embedded and how the response field is populated.
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-05.