← all guides

Header order and what it gives away

Most people troubleshooting a blocked scraper start with the IP. Is it a datacenter range, is it flagged, does it need to rotate faster. Fewer people look at the request itself, and that’s usually the bigger tell. Every HTTP client, whether it’s Chrome, Firefox, curl, or a Python script built on requests, sends its headers in a specific order with specific casing. That order is boring, consistent, and almost never touched by the person writing the scraper, which is exactly why it’s useful to whoever is trying to spot one.

Headers have an order because software has habits

A browser doesn’t generate its headers randomly each time. The rendering engine builds the request through a fixed code path, and that path appends headers in the same sequence every single time, for every request, on every page load. Chrome’s header order looks different from Firefox’s, which looks different from Safari’s, and each one is remarkably stable across versions unless something in the networking stack actually changes.

HTTP/1.1 headers are plain text lines, so the order they appear on the wire is just the order the client wrote them. HTTP/2 and HTTP/3 use compressed, indexed formats (HPACK and QPACK), and technically headers can be grouped differently, but the pseudo-headers (:method, :authority, :scheme, :path) still show up in a fixed order per client, and the regular header list that follows still reflects the sending library’s internal order. The mechanism changed, the fingerprint didn’t go away.

This is also why casing matters. HTTP header names are case-insensitive per spec, servers are supposed to treat User-Agent and user-agent as identical. In practice almost nothing normalizes them at the point of comparison, because doing so would throw away a useful signal. Real browsers have consistent, well-known casing conventions. A lot of HTTP libraries default to their own casing, often all lowercase, because that’s simpler to implement and the spec allows it. A server that logs raw header casing can tell the difference between “a browser” and “a library pretending to be a browser” in a single glance, before it even looks at the values.

What a default HTTP client gets wrong

If you write a scraper with requests, axios, node-fetch, or similar and set nothing but a User-Agent, you get a header set that no real browser would ever send. A few concrete gaps show up over and over:

  • Missing client hints. Modern Chromium sends sec-ch-ua, sec-ch-ua-mobile, and sec-ch-ua-platform on most requests. A client claiming to be Chrome 120 without any of these looks like it’s lying about its own identity.
  • Missing fetch metadata. Real browser requests carry sec-fetch-site, sec-fetch-mode, sec-fetch-dest, and often sec-fetch-user. These describe the context of the request (top-level navigation vs. a subresource fetch vs. a cross-site request), and a bare scraper doesn’t generate them because it doesn’t have a rendering context to describe.
  • A short, generic Accept header. Browsers send a long, specific Accept value listing preferred MIME types with quality weights. Default HTTP libraries often send */* or a minimal list.
  • No Accept-Language, or one that doesn’t match anything. Browsers reflect actual OS/browser locale settings. A scraper either omits it or hardcodes something that never changes across requests, which is its own pattern.
  • Header order that doesn’t match the claimed User-Agent. This is the one people miss even after fixing the values above. You can set every header a real Chrome request would set and still get flagged, because your library appended them in its own internal order rather than Chrome’s.

None of this is exotic. It’s the direct, visible consequence of writing HTTP requests with a general-purpose library instead of a browser engine.

How this actually gets used against a request

Detection vendors (Akamai, Cloudflare’s bot management, HUMAN, DataDome, and others) don’t rely on header order alone, and it’s worth being precise about that instead of overstating it. Header order and casing are one signal among several that get combined into a client fingerprint. The other major piece is the TLS handshake itself: the cipher suite list, extension order, and supported curves a client offers during the TLS ClientHello form a separate fingerprint (commonly referred to as JA3 or JA4) that’s independent of anything in the HTTP layer. A request can have perfect HTTP headers and still get flagged because the TLS handshake underneath it came from OpenSSL’s default ordering instead of a real browser’s.

What these systems are actually doing is checking for internal consistency. A request claiming to be Chrome 120 on Windows should have a TLS fingerprint that matches known Chrome 120 handshakes, an HTTP/2 pseudo-header order that matches Chromium’s, an HTTP header order and casing that matches Chromium’s, and a full set of client hint and fetch metadata headers. When some of those line up and others don’t, the mismatch itself is the signal, independent of anything about the IP the request came from. This is why the framing “get a good proxy and you’re fine” is wrong, and worth saying plainly: proxy quality solves an IP reputation problem, not a client identity problem. They’re different layers, checked by different parts of the same detection pipeline.

Why rotating IPs doesn’t fix a bad fingerprint

This is the part that catches people running scrapers behind residential or mobile proxies. If every request through a rotating IP pool carries the exact same header order and TLS fingerprint, the detection system can link those requests together as the same client regardless of which IP each one came from. The IP rotation solves rate limiting and IP-based reputation scoring. It does nothing for a fingerprint that says “this is a Python script” on every single request. In some setups, a stable, distinctive fingerprint moving across hundreds of IPs is a stronger signal of automation than a single IP making repeated requests, because real users don’t share a browser fingerprint across a residential IP pool.

What staying clean actually looks like

A compliant scraper that wants to look like ordinary traffic needs internal consistency between three things: the claimed User-Agent, the TLS handshake, and the HTTP header set including order and casing. In practice this usually means using a client library or tool built specifically to reproduce a real browser’s network stack, rather than trying to hand-assemble the right headers on top of a generic HTTP client. Headless browser automation (real Chromium or Firefox instances) sidesteps most of this by construction, since the requests come from an actual browser engine. Purpose-built HTTP clients that replicate specific browser TLS and header fingerprints exist for the same reason.

None of this makes a scraper undetectable, and it shouldn’t be sold that way. Detection systems also look at request timing, mouse and scroll behavior on rendered pages, cookie and session continuity, and outright behavioral patterns that no amount of header tuning fixes. A consistent fingerprint reduces one specific class of flag; it doesn’t override the rest of a site’s defenses, and it doesn’t change what’s appropriate to scrape in the first place. Sites set rate limits and terms for reasons, and matching a browser’s header profile is about not looking like an obviously broken client, not about getting around rules a site has put in place on purpose.

If you’re debugging why a scraper gets blocked even through clean residential IPs, check the header order and TLS fingerprint before touching the proxy pool. It’s often the actual cause, and it’s the kind of thing that’s easy to overlook because the headers “look right” when you eyeball the values instead of the order they arrive in.

If you want more breakdowns like this on how scraping infrastructure actually behaves under detection, and honest comparisons of the proxy types behind it, come find us Proxy Scraping.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →