← all guides

AWS WAF Bot Control and what it flags in scrapers

If you scrape retail, travel, ticketing or SaaS pricing pages long enough, you will hit a site sitting behind AWS WAF Bot Control. It is easy to miss at first because it does not announce itself the way Cloudflare or DataDome do. There is no branded interstitial, no “checking your browser” page with a logo. You get a 403 with a short body, or a 202 with a blob of JavaScript, or a 405 on what should have been a plain XHR. Then your success rate drops from 98 percent to 40 and your logs are full of responses that look like ordinary errors.

The stakes are practical. Bot Control is sold by AWS as a managed rule group that any customer can switch on in a few clicks, which means the sites using it range from careful enterprise teams to a small shop that ticked a box and never looked again. The first group has tuned rules, token handling and rate limits. The second usually runs the defaults, and the defaults are quite predictable. Knowing which is which changes how much effort you should spend.

This article is the deep-dive I wish I had when I first started debugging these blocks. I will cover what the rule group actually inspects, how the token and challenge flow works, what I have seen trip scrapers in production, and how I test for it without guessing. I write this from the operator side, running scrapers out of Singapore against targets in the US, EU and Southeast Asia. I do not have access to anyone’s WAF config, so where I describe internals I stick to AWS’s public docs and to what I can observe from the outside.

Background and prior art

AWS WAF has existed since 2015 as a rules engine: you write conditions on IPs, headers, URI strings and rate, and it allows, blocks or counts. Bot Control arrived later as an AWS Managed Rule group (AWSManagedRulesBotControlRuleSet) and has two inspection levels, Common and Targeted. AWS documents both in the AWS WAF Bot Control guide, and the individual rules and labels are listed in the Bot Control rule group reference. Those two pages are the primary sources for almost everything technical below.

The design choice worth understanding is that Bot Control does not decide anything on its own the way a standalone bot vendor does. It attaches labels to requests, and the customer’s web ACL decides what to do with them. A rule can label a request bot:category:http_library and the site owner might block it, count it, or ignore it. That is why two sites using “AWS WAF Bot Control” can behave completely differently against the same scraper. The rule group is a sensor. The web ACL is the policy. Most of the pain in scraping these sites comes from that split, because you are never fighting one system, you are fighting one sensor plus whatever the owner wired to it.

For prior art on the scraper side, most of what is written treats WAF blocking as an IP reputation problem. With Bot Control that is only about a third of the picture. The rest is user agent and header consistency, TLS and browser fingerprinting via the token, and behaviour over time. If you have read my earlier piece on header order and what it gives away, the Common level is largely a productised version of those checks.

The core mechanism

Common level: static inspection

Common inspects each request in isolation. It does not need JavaScript on the client, and it does not need session history. It looks at things like the user agent string, headers, and the source IP against known data center ranges, then assigns labels. The categories AWS documents include ones such as CategoryHttpLibrary, CategoryScrapingFramework, CategorySearchEngine, CategorySeo, CategorySecurity, CategoryMonitoring, CategoryContentFetcher and CategoryMiscellaneous, plus signal rules like SignalNonBrowserUserAgent and SignalKnownBotDataCenter. AWS adds and renames rules over time, so check the current list rather than trusting my memory of it.

What that means in practice:

  • a default python-requests/2.x, curl/8.x, Go-http-client/1.1 or okhttp user agent gets labelled as an HTTP library and is often blocked outright
  • a user agent naming a known scraping framework gets the scraping framework label
  • a browser-looking user agent with missing or oddly ordered headers can trip the non-browser signal
  • traffic from well-known cloud ranges can pick up the data center signal, which is the reason datacenter proxies get flagged so quickly on these sites
  • verified bots such as legitimate search engine crawlers are validated separately, so spoofing a Googlebot user agent from a random IP fails the verification and gets you labelled as an unverified bot

The last point catches people. Pretending to be a crawler is worse than being an anonymous browser, because AWS checks the claim.

Targeted level: tokens, fingerprints and behaviour

Targeted is where scrapers that pass Common start to fail. It adds detection that depends on a client-side component. The customer integrates the AWS WAF JavaScript SDK (or uses the challenge and CAPTCHA actions), which runs in the browser, collects signals, and produces a token stored in the aws-waf-token cookie. The docs cover this flow in CAPTCHA and Challenge in AWS WAF.

The token is the centre of the mechanism, so it is worth being precise about it:

  • it is issued after the client completes a challenge, which is a silent JavaScript proof-of-work and environment check for the Challenge action, or an interactive puzzle for the CAPTCHA action
  • it carries an expiry, controlled by the “immunity time” configured on the web ACL, which defaults to 300 seconds
  • it is tied to the client characteristics the WAF observed at issue time, and Targeted rules watch for the same token appearing from multiple IPs, countries, ASNs or TLS fingerprints
  • requests that arrive with no token at all, on paths that should have one, are themselves a signal

The Targeted rules I see named most often in labels and logs are TGT_VolumetricSession, TGT_TokenReuseIp, TGT_TokenReuseCountry, TGT_TokenReuseAsn, TGT_TokenReuseTlsFingerprint, TGT_SignalAutomatedBrowser, TGT_SignalBrowserInconsistency and a family of TGT_ML_CoordinatedActivity rules that use machine learning across many requests to spot coordination. Again, AWS maintains the authoritative list.

The important shift is that Targeted evaluates sessions, not requests. One session making 600 requests in five minutes gets a volumetric flag that a hundred sessions making six requests each do not. That is why rotating proxies alone stops helping. If you rotate the IP but keep the token, you trigger token reuse. If you rotate the token but keep the IP pool tiny, you trigger the ML coordination rules.

Actions: block, count, challenge, CAPTCHA

What you see depends on what the owner attached to those labels:

  • Block returns a 403, which is what you see on Common-level hits
  • Challenge on a page load returns an interstitial that runs JavaScript, then reloads with a token. On non-navigation requests, the response is a 202 or similar status with a challenge payload rather than your data
  • CAPTCHA returns a puzzle page, and for non-page requests such as XHR it returns a 405, which is a strange status that confuses parsers written to only expect 200, 403 and 429
  • Count lets the request through and only adds the label, so a site can be watching you long before it blocks you

That last one is worth remembering. If your success rate is fine but you suddenly start failing a week later, the owner may have moved from Count to Block after reviewing their own data.

Cost as a hidden constraint

Bot Control is billed per request inspected, and Targeted costs noticeably more than Common. Pricing is on the AWS WAF pricing page and it changes, so I will not quote numbers I cannot verify today. The operational consequence is what matters: many owners scope Bot Control to a subset of paths using a scope-down statement, for example only /login, /checkout and /api/*, because they do not want to pay to inspect every image request. This means the same site can be wide open on product pages and hostile on search or cart endpoints. I always test path by path before assuming a whole domain is protected or unprotected.

Worked examples

These are illustrative setups built from the patterns I see repeatedly, with the numbers chosen to show the arithmetic. They are not measurements from a specific named retailer.

Example 1: a plain HTTP client against a Common-only site

Setup: a Python script using requests with default headers, 20 requests per minute, through a datacenter proxy pool of 50 IPs on a well-known cloud provider.

What happens: the user agent triggers the HTTP library category and the source ranges trigger the data center signal. Both are Common-level, so no JavaScript is needed to catch it. Result is a 403 on nearly every request, and the fix is nothing to do with rate.

The fix in steps:

  • send a real, current browser user agent, and make the rest of the headers agree with it (Accept, Accept-Language, Sec-Fetch-*, and Sec-CH-UA if the UA claims Chromium)
  • match header order to the browser, since some HTTP libraries send headers in an order no browser uses
  • move to residential or mobile IPs if the data center signal is what is being counted

A minimal header set that gets past Common on the sites I have tested looks like this:

import httpx

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
                  "(KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
    "Accept-Language": "en-US,en;q=0.9",
    "Accept-Encoding": "gzip, deflate, br",
    "Upgrade-Insecure-Requests": "1",
    "Sec-Fetch-Dest": "document",
    "Sec-Fetch-Mode": "navigate",
    "Sec-Fetch-Site": "none",
    "Sec-Fetch-User": "?1",
}

with httpx.Client(http2=True, headers=headers, proxy="http://user:[email protected]:8000") as c:
    r = c.get("https://target.example/product/123")
    print(r.status_code, r.headers.get("x-amzn-waf-action"))

The x-amzn-waf-action response header is one I look for. When present it tells you the WAF acted on the request, which is a quick way to separate a WAF block from an origin 403. Not every configuration exposes it, so its absence proves nothing.

Numbers: at 20 requests per minute over 50 IPs, each IP sees about one request every 2.5 minutes. Rate was never the issue. If you had spent a day slowing the scraper down, you would have learned nothing.

Setup: Playwright with Chromium, one browser context per job, 8 workers, residential proxies rotating on every request. Target has Targeted level on /search with a Challenge action.

What happens: the first request to /search runs the challenge and receives an aws-waf-token. The next request goes out through a different residential IP but reuses the cookie. The token was issued to IP A and is now presented from IP B in another ASN. That is TGT_TokenReuseIp, and probably TGT_TokenReuseAsn too. The site re-challenges or blocks, and because the browser is re-solving a challenge on almost every request, the volume of challenge solves looks like automation by itself.

The fix:

  • pin an IP to a browser context for the life of the token, meaning a sticky session of at least the immunity time (300 seconds by default, check the site’s actual setting by watching when the token stops working)
  • reuse the context and token for many requests rather than launching fresh ones
  • keep the TLS fingerprint constant within a session, which browsers do naturally but HTTP clients that share a token with a browser do not

Numbers: 8 workers each holding a sticky IP for 10 minutes and making a search every 20 seconds is 30 searches per worker per session, 240 per 10-minute window across the fleet, and 8 challenge solves instead of 240. The solve count dropped by 97 percent from the arithmetic alone, and that is before considering the reuse flags. I wrote about the proxy setting side of this in Playwright proxy settings that silently do nothing, and a misconfigured proxy in the launch options is one reason a token seems to be “reused” from a different IP when you thought it was pinned.

Example 3: hybrid, browser for the token and HTTP client for the volume

Setup: the product data comes from a JSON endpoint (/api/products?page=N) under Targeted, and rendering each page in a browser costs too much bandwidth.

Approach: a browser context loads one HTML page through a sticky IP, receives the token, and the script reads the cookies. A lightweight client then makes the JSON calls from the same IP with the same cookie for the token’s remaining life.

The trap: the HTTP client’s TLS fingerprint differs from Chromium’s. If the site checks TGT_TokenReuseTlsFingerprint, the token issued to a Chromium handshake is now being used by a different one. Libraries that mimic browser TLS handshakes (curl-impersonate and the Python bindings built on it) exist for this reason. I have had mixed results, so I test it on each target instead of assuming it will hold. If it fails, do the JSON calls from inside the page context with fetch, which keeps the browser’s TLS stack and costs a little more CPU.

Edge cases and failure modes

Verified bot spoofing backfires

It is tempting to send a Googlebot or Bingbot user agent because sites usually want those crawlers. AWS validates verified bots, so an unverified request claiming to be one is labelled as a spoof and is handled worse than an ordinary unknown client. Counter: do not claim to be a crawler. If a site genuinely allows an API or feed for your use, use that instead. Also read the target’s robots.txt rules and terms before you scrape, since those are the rules that bind you regardless of what the WAF lets through. This is not legal advice.

The 405 and 202 that look like success or noise

CAPTCHA on an XHR path returns 405, and Challenge can return a 2xx-family status with a challenge body. A scraper that treats any 2xx as success will happily parse the challenge JavaScript as data and store garbage. Counter: validate the response body, not the status. Check for an expected key or selector, and treat the presence of awswaf script references or the x-amzn-waf-action header as a block. My notes on alerting on block rate instead of errors cover why I count content-validation failures as blocks and not as successes.

Token expiry mid-job

The token has a lifetime. When it lapses on a long job, the next request gets re-challenged, and a client without a JavaScript engine cannot answer. The failure shows up as a step change in error rate at a fixed interval, often exactly 5 minutes, which is a giveaway. Counter: track token age per session, refresh before expiry by loading a page in the browser context, and retire sessions cleanly rather than letting them fail. If you see failures at a suspiciously round interval, suspect immunity time before you suspect IP quality.

Retries that feed the coordination model

When a request fails, a naive retry loop sends it again immediately, often from a new IP with the old cookie. That is a perfect token reuse pattern and it also inflates volume. Counter: back off, keep retries on the same session or drop the token with the session, and cap retries per URL. I go through the mechanics in handling retries without amplifying your own block rate.

Path-by-path inconsistency

Because owners scope Bot Control to certain paths, your test on the homepage tells you nothing about /api/ or /cart. Counter: probe each path family with a clean session and record status, headers and body signature separately. I keep a small table per target: path pattern, WAF header present or not, action seen, and whether a token was required.

What we learned in production

The biggest lesson is that Common and Targeted need to be diagnosed separately, because the fixes are different. If a plain browser-like request from a residential IP works and your scraper does not, the difference is Common-level: headers, user agent, TLS or IP class. If a real browser works for a few requests and then degrades, it is Targeted: token reuse, session volume or coordination. I run every new target through a fixed sequence. First a single request from my own machine in a normal browser, then the same via curl with a full header set, then through the proxy, then through Playwright. The step where it starts failing tells me which layer I am fighting. It takes about 15 minutes and saves days of tuning the wrong knob.

The second lesson is that sticky sessions matter more than pool size on these targets. A pool of 200 residential IPs rotated per request did worse for me than 40 IPs held for the length of a token. Bandwidth is the cost, since holding a browser open and loading a page for the token is heavier than a bare request, so I budget for it. My rough rule is one browser page load per token life, then everything else through the cheapest client that survives the TLS check. And when a target moves from Count to Block with no change on my side, I treat it as a signal to slow down and reassess, not to push harder. If you are running many identities in parallel on other platforms, the same session-versus-request thinking applies, and the multiaccountops.com blog covers the session isolation side of that in more depth, while antidetectreview.org covers browser fingerprint tooling.

For more on the proxy side and the rest of the troubleshooting set, the blog index has the full list.

References and further reading

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-29.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →