← all guides

Soft blocks that return two hundred

The status code lies

Every scraper starts out trusting the HTTP status code. A 200 means it worked, a 403 means it didn’t, a 429 means slow down. That mental model holds up fine until you start scraping sites that have invested real engineering time into bot detection, because a lot of those sites stopped using error codes as their primary defense years ago.

A soft block is a response that comes back with a 200 status and looks, structurally, like a normal page, but doesn’t contain the data you asked for. No error, no captcha challenge, no rate-limit header. Just a page that renders, parses, and passes every check your scraper runs on the transport layer, while quietly withholding the thing you actually wanted.

If you’re only checking response.status_code == 200 before moving on, a soft block will slip straight through your pipeline and you won’t find out until someone looks at the output and asks why half your product listings are empty or why your price field has been null for three days.

Why a site would rather serve 200 than 403

From the defender’s side, a 403 or a captcha wall is a confession. It tells the requester, and anyone monitoring the requester’s error rate, exactly when detection triggered. That’s useful information for the person on the other end, because it lets them adjust whatever behavior tripped the block and try again with a cleaner request.

A 200 with degraded content gives away nothing. The scraper doesn’t get a clean signal to react to. If the operator running it isn’t validating content shape, the pipeline just keeps running and keeps collecting garbage, which from the site’s perspective is a much better outcome than an obvious block that gets diagnosed and adjusted around in an afternoon. It also avoids tipping off legitimate users who might be caught in a false positive, since a real visitor and a flagged one both see a 200 and neither gets an alarming error page.

This is standard practice on sites that take scraping seriously: e-commerce platforms that don’t want competitors price-matching in real time, ticketing and travel sites managing inventory visibility, and social platforms limiting what an unauthenticated or suspicious session can see. None of this is a secret technique. It’s a well documented category of anti-bot response, and understanding how it’s built is what lets a scraper operator run cleanly instead of blind.

What actually gets swapped in

The mechanics vary by site but they cluster into a few recognizable patterns:

Empty or partial data sets. The listing page loads, the layout is intact, but the results array is empty or truncated to a handful of items regardless of what the query should return.

Stale or cached content. The page serves a snapshot from before the block triggered, so numbers stop moving. A price page that hasn’t changed in six hours despite a volatile market is a strong tell.

Decoy or sanitized content. Some sites serve a version of the page with certain fields stripped or replaced with placeholder values, close enough to pass a shallow check but useless for anything downstream.

A different render path entirely. Server-rendered HTML that normally carries the data in the initial payload gets swapped for a client-side shell that expects JavaScript to fetch the real content from an API the flagged session can’t reach. The scraper gets a 200 and a page full of <div id="root"></div>.

Silent redirection to a lookalike page. The URL and status look right but the content is a generic category page or a “no results” template dressed up to resemble a real one.

None of these trip a status-code check. All of them are visible if you look at the actual bytes.

How this differs across proxy types

Soft blocks are a content-layer defense, but the proxy layer still affects how often you run into them. Detection systems that use behavioral and reputation scoring, not just IP category, weigh things like ASN, how many other sessions have recently come from the same address, and how consistent the request pattern looks over time.

Datacenter IPs tend to get flagged faster because their ASNs are well known and easy to bucket, so a scraper running on datacenter infrastructure will often see soft blocks appear earlier in a session, sometimes on the first few requests to a sensitive endpoint. Residential IPs carry better reputation because they sit in consumer ASNs alongside real traffic, but a residential IP that’s been reused by many scrapers before you inherits whatever reputation damage they left behind, which is one reason pool quality and churn matter as much as the category label. Mobile IPs behind carrier-grade NAT get a different kind of leniency because blocking one IP risks blocking hundreds of unrelated phones on the same carrier, but that leniency isn’t the same as invisibility, and a session that behaves like a scraper on a mobile IP will still get soft-blocked, just possibly with more patience from the detector first.

None of this makes any proxy type immune. It shifts where in the request pattern the soft block tends to show up, and that’s useful information for diagnosing what’s happening, not a guarantee about what will happen.

Building detection for it, not around it

The fix isn’t a trick against the site, it’s better validation inside your own pipeline. A scraper that only checks the status code is flying blind. One that checks the shape of the response can catch a soft block the same request cycle it happens, which matters because reacting immediately is what keeps a flagged session from generating hundreds of useless requests before anyone notices.

Practical checks that catch this category of block:

  • Compare response size and structure against a known-good baseline for that page type, and flag outliers.
  • Assert on the presence of specific data fields you expect, not just that the HTML parsed without throwing.
  • Track whether the same values are repeating across requests that should return different data. Static numbers on a page that should be live is a strong signal.
  • Diff a sample of responses against a manually verified page fetched through a clean, unflagged path, and look for structural drift over time.
  • Watch for a sudden shift from server-rendered content to a JS shell on a site that previously served data inline.

None of this is about defeating a captcha or forging a fingerprint. It’s the same discipline as monitoring any pipeline for silent failure: assume the transport layer can lie to you and verify the payload.

What a clean response looks like

When a scraper detects a soft block, the compliant move is to back off, not push harder. That means slowing the request rate on that target, rotating to a different, cleaner IP rather than hammering the same one, and respecting the site’s published terms of service and robots.txt for that endpoint. If a site has made clear through its terms that a particular kind of automated access isn’t permitted, a soft block that starts showing up is a signal to stop, not a puzzle to solve.

There’s no proxy or rotation strategy that removes this risk entirely, and any provider or technique claiming otherwise is describing marketing, not a mechanism. What you can build is a scraper that fails loudly to you, even when the site fails quietly to it, so a soft block turns into a fixable data quality issue in your own logs instead of a mystery three weeks of bad numbers later.

If you want more of this kind of teardown, on the mechanics of detection and what running proxy infrastructure honestly looks like from the inside, you can find the rest of it on the Proxy Scraping home page.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →