← all guides

Debugging a scraper that only fails in production

The bug that isn’t in your code

Every scraper operator has lived this: the script runs fine on a laptop, passes every test, then gets deployed to a server and starts throwing 403s, CAPTCHAs, or silently empty responses within the hour. The instinct is to reread the parsing logic, check for a typo in a selector, add another try/except block. Nine times out of ten, none of that is the problem. The code is identical. What changed is everything around the code: the network path, the IP, the concurrency, the headers a real browser would have sent but your headless client didn’t.

Scraping is one of the few kinds of software where the environment is part of the program. A normal API client behaves the same whether it runs from a laptop or a data center, because the server on the other end doesn’t care where the request came from. A scraper hits a target that is actively trying to tell “a person browsing” apart from “a script fetching pages,” and the signals it uses for that are mostly environmental: IP reputation, TLS handshake shape, request timing, header order, and behavior over a session. Local testing almost never reproduces those signals honestly, which is why “works on my machine” is close to meaningless for a scraper.

Start with the egress, not the parser

The single most common cause of prod-only failures is a change in egress IP. A laptop on a home connection usually sends traffic from a residential ISP block. A server, especially a cloud VM, sends traffic from a data center ASN. Sites that care about scraping maintain IP reputation data at the ASN and subnet level, and data center ranges are cheap to flag wholesale because almost no real consumer traffic originates there. Your code didn’t get worse. Your IP got worse.

This is easy to confirm and often skipped. Run the exact same request, same headers, same code path, from both environments and compare status codes and response bodies directly, not just “it worked” versus “it didn’t.” If the local run succeeds and the prod run gets a 403, a CAPTCHA redirect, or a suspiciously generic 200 with no real data in it, you’ve isolated the variable to the network identity, not the logic.

If you’re already routing production traffic through proxies, the debugging question shifts: which proxy, and what’s its reputation history. A shared datacenter IP that’s been hammered by other tenants before you touched it will behave differently than a fresh one, and a proxy pool with no health monitoring will happily keep routing traffic through IPs that are already flagged. This is a maintenance problem as much as a coding one. A pool needs some way to track failure rates per IP and retire the ones that are consistently getting blocked, the same way you’d take a bad server out of a load balancer rotation.

The fingerprint gap between local and headless

If your local testing was done with a normal browser, or even manually clicking through the site, and production runs headless, you’re comparing two different clients, not two environments running the same client. Headless browsers and HTTP libraries have fingerprints: TLS handshake parameters (cipher order, extensions), HTTP header order and casing, missing headers a real browser always sends, and JavaScript environment quirks like navigator properties that don’t match a real device. None of this is exotic. It’s just detail that’s easy to get right by accident in a real browser and easy to get wrong in a script.

The fix isn’t a trick, it’s parity. Capture the actual request your production client sends, headers and TLS parameters included, and diff it against what a real browser sends to the same endpoint. Tools that dump raw TLS ClientHello data or full request headers make this a five-minute comparison instead of a guessing game. If your scraper is missing an Accept-Language header, sends headers in an order no browser uses, or negotiates a TLS cipher suite your library defaults to but Chrome never would, that’s a concrete, fixable gap, and it’s usually a bigger factor than which specific proxy you’re using.

Concurrency changes the traffic shape

Local testing tends to be one request, watch it work, done. Production runs the same scraper across dozens or hundreds of concurrent workers, often against the same handful of target domains. That changes the traffic pattern from “a person browsing” to “a burst from one source,” and rate-based detection is built exactly to catch that shape. A site doesn’t need to fingerprint your client at all if fifty requests a second are arriving from the same subnet with identical timing.

This is worth checking with real numbers before assuming it’s an IP quality problem. Log request volume per IP per minute in production and look at whether failures cluster around the moments concurrency spikes. If a single proxy is absorbing a disproportionate share of your worker pool’s traffic, that’s a routing and pool-sizing issue, not something a “better” proxy fixes on its own. Sane concurrency limits per IP, and enough IPs in rotation to keep any single one from looking like a traffic spike, are basic hygiene for anything running at scale, and they’re the difference between traffic that looks organic and traffic that looks like a bot farm because, structurally, it is one.

Infrastructure noise that looks like blocking

Not every prod-only failure is detection at all. A few boring causes show up constantly:

DNS resolution differs between environments, especially if production runs in a container with a different resolver configuration, and a scraper that resolves to a stale or wrong IP will fail in ways that look like blocking but aren’t.

Clock skew on a production host can break TLS handshakes or session tokens that depend on timestamps, producing failures that have nothing to do with the target site’s bot detection.

Memory pressure under concurrent load can truncate responses or cause a scraper to parse a partial page as if it were complete, which then gets misread as “the site changed its markup” when actually the connection just got cut early.

Retry and backoff logic that was never exercised locally, because you only ever ran one request at a time, can behave badly under real failure rates and turn a handful of legitimate rate limits into a much larger blocked window.

None of these are glamorous, but they’re often faster to rule out than chasing a proxy problem, and ruling them out first saves you from “fixing” something that was never broken.

A debugging order that actually narrows it down

Change one variable at a time. Run the identical code from the production host manually and compare it to the same run from local. Check IP reputation and ASN type for whatever egress you’re actually using in each environment. Diff the raw request, headers and TLS included, between your scraper and a real browser hitting the same page. Log status code and response size distributions per proxy or per IP, not just an aggregate success rate, so a handful of bad IPs don’t hide inside a good average. Reproduce production concurrency locally against a low-stakes endpoint before assuming it’s the target site’s defenses and not your own traffic shape.

A scraper that only fails in production is almost always telling you something true: production is not the same client the target site sees when it’s happy. Sites that care about distinguishing bots from people are reading exactly the signals that differ between a laptop and a fleet of concurrent cloud workers, IP reputation, request fingerprint, and traffic shape. None of this means a scraper can be made undetectable or guaranteed to run cleanly forever. It means the failure has a cause, and the cause is usually visible once you stop looking at the parser and start looking at what actually left the machine.

If you’re trying to figure out which of these is hitting your own stack, or want an honest comparison of proxy types before you build a production pool around one, that’s what we write about.

Read more on Proxy Scraping

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →