← all guides

Two identical requests, two different answers

You send the same GET request twice, thirty seconds apart. Same URL, same headers, same everything you control. The first response has the product price you expect. The second one has a different price, or a login wall, or a page that looks almost right but is missing half the listings. Nothing in your code changed. This is one of the most common things that trips up people running scrapers at scale, and it usually isn’t a bug in your scraper. It’s a sign that “identical request” was never as identical as it looked.

Your request isn’t as identical as you think

From your terminal, two requests to the same URL look the same. From the server’s side, they rarely are. A request carries more than the URL and headers you typed. It carries the IP address it came from, the TLS handshake that opened the connection, the exact order and casing of headers your HTTP client sent, and whatever cookies or session tokens got attached along the way. Change any one of those and you’ve changed the request, even if your code thinks it sent the same thing twice.

If you’re running requests through a rotating proxy pool, the IP is almost guaranteed to change between calls, sometimes on every call. A new IP means a new geographic origin as far as the target server is concerned, a different reputation history, and possibly a different network path entirely. Two requests that look byte-identical in your logs can arrive at the origin server looking like they came from two different people in two different countries.

Server-side reasons the answer changes

Most inconsistency has nothing to do with anti-bot systems and everything to do with how modern web infrastructure is built.

Load balancing. Large sites don’t run on one server. A request gets routed to whichever backend is free, and if those backends aren’t perfectly synced (a cache that hasn’t refreshed on one node, a database replica that’s a few seconds behind), you’ll get slightly different data depending on which one answered. This is completely normal and has nothing to do with you.

CDN edge caching. If you’re hitting a page served from a CDN, your response depends on which edge location answered and what that edge had cached. A proxy that puts you in Singapore for one request and Frankfurt for the next will likely hit two different edge caches with two different snapshots of the page, especially for content that updates frequently like prices or stock levels.

Personalization and A/B testing. A lot of sites vary content by cookie, by session, or by a random bucket assigned on first visit, testing layouts, pricing displays, or feature rollouts against different slices of traffic. If your requests don’t carry a consistent session, you can get bucketed differently each time, and the “difference” you’re seeing is the site working exactly as designed.

Geo-targeting. E-commerce and travel sites routinely serve different prices, currencies, or inventory by detected location. If your proxy rotates you across regions, you’re not seeing inconsistency in their system, you’re seeing their system correctly doing its job on inconsistent input.

Soft throttling. Some servers respond to request volume by quietly degrading what they return rather than issuing a hard block, serving a cached or simplified version of a page once a client crosses some internal rate threshold. That behavior looks like flakiness from the outside but is a deliberate load-management decision on their end.

Where bot detection fits into this

Anti-bot and WAF systems are one more source of response variance, and it’s worth understanding how they actually work rather than treating them as a mystery. Most modern detection systems don’t make a binary allow-or-block decision per request. They score a session on signals accumulated over time: IP reputation, TLS and HTTP fingerprint consistency, request timing, mouse or scroll behavior on JS-rendered pages, and whether the traffic pattern looks like a browser or a script. Based on that score, the system can choose from a range of responses, not just “200 OK” or “403 Forbidden.” It might serve a cached, stripped-down, or intentionally stale version of a page to traffic it doesn’t fully trust, hold a request behind a JS challenge, or quietly rate-limit without saying so.

This is defensive infrastructure working as intended, and it’s a legitimate reason two calls that look the same to you get treated differently by the server. It is not something to route around. If a site is serving you degraded content because your traffic pattern reads as automated, the correct response as an operator is to look at what your traffic pattern is actually doing, slow it down, and bring your request behavior in line with what the site’s terms of service allow, not to search for a trick to make the detection stop noticing. Detection logic like this changes constantly and unpredictably on the target’s side, and any specific workaround is a short-lived arms race you don’t want to be running, on top of the legal and ethical line it puts you on.

Proxy type changes the shape of the inconsistency

The kind of proxy behind your scraper affects what kind of inconsistency you should expect.

Datacenter proxies sit on IP ranges that are well documented and easy for a target to flag as hosting infrastructure rather than a residential connection. If a target treats datacenter ranges more cautiously, you’ll see more soft-degraded or challenge responses from that pool, and the inconsistency will correlate strongly with which IP block you landed on.

Residential proxies route through real consumer ISP connections, which generally read as more trustworthy, but that trust comes with less control. You’re at the mercy of whatever that household’s connection quality and history actually is, and a residential IP that a previous user hammered with abusive traffic carries that reputation with it whether you caused it or not.

Mobile proxies exit through carrier networks, which often share a small pool of public IPs across large numbers of subscribers through carrier-grade NAT. That means your traffic gets mixed in with real phone traffic on the same address, which can help trust signals, but it also means the same IP might be serving requests from many unrelated devices at once, and carrier IPs get reassigned constantly. You can query the same “IP” twice and actually be talking to a completely different phone’s network path.

None of these characteristics make a proxy type good or bad in the abstract. They change what kind of consistency you can expect and what you need to log to make sense of what you’re seeing.

Building a scraper that expects this

The practical fix isn’t to chase perfect consistency, it’s to stop assuming it exists and design around its absence.

Log the full context with every response, not just the payload: which proxy IP handled the request, what geo and ASN that IP resolved to, response status, response size, and a hash of the body. When something looks off, you want to be able to ask “did this vary because of geo, because of caching, because of a session mismatch, or because of something else” and actually answer it from your logs instead of guessing.

Treat a single sample as unreliable for anything you plan to act on. If a price or a stock count matters, pull it a few times across different proxy exits before trusting it, and flag results that disagree rather than silently averaging them.

Fingerprint known block or challenge pages by their size, structure, or specific markers, and separate “this is a legitimate alternate version of the page” from “this request got flagged.” Conflating the two is where most debugging time gets wasted.

Keep session state coherent within a single scraping run. If you rotate the IP mid-session while keeping the same cookies, you can trip the exact personalization and consistency mismatches described above for reasons that have nothing to do with detection.

And build in backoff, not brute force. If a target’s responses are degrading under your request volume, that’s information, not an obstacle to push through faster.

Choosing a provider with this in mind

When you’re evaluating a proxy provider for scraping, don’t take “residential” or “unlimited rotation” at face value. Test the specific thing you need: how sticky can a session stay on one exit IP, how tightly can you target a geo, how large is the actual pool behind the marketing number, and what happens to your response consistency when you hold one IP for ten minutes versus rotating every request. Those are things you can verify yourself against your own target sites, and the answer will differ by provider, by target, and by day. Anyone telling you a proxy is undetectable or guarantees clean results hasn’t run it against a real production target long enough to find out otherwise.

If you want more of this kind of breakdown on how proxies, rotation, and scraping infrastructure actually behave in production, you can find the rest of our writing and videos at Proxy Scraping.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →