Why your scraper alerts should watch block rate, not error rate
Most scraping setups alert on the wrong thing. Someone wires up a check for non-2xx status codes, calls it monitoring, and moves on. Then a target site starts serving CAPTCHA pages with a 200 status, the error dashboard stays green, and nobody notices the pipeline has been collecting garbage for three days. If you run scrapers at any real scale, this happens to everyone eventually. The fix isn’t a smarter error alert. It’s a different metric entirely: block rate.
Errors and blocks are not the same thing
An error is a status code your HTTP client recognizes as a failure: a 403, a 429, a connection timeout, a DNS failure. These are useful to track, but they only catch the blunt forms of blocking. A site that wants to shut out scrapers without tipping its hand rarely does it that bluntly.
What actually happens more often is that the request succeeds at the transport level. The server returns 200. The body just isn’t what you asked for. It might be a CAPTCHA challenge page, a “verify you are human” interstitial, a JavaScript challenge that never resolves without a real browser, or a thinned-out version of the page missing the data you’re after. From the outside, your scraper looks healthy: response received, status ok, latency normal. From the inside, the run is producing nothing usable.
This is the gap that block rate is meant to close. It measures how often you’re being served a block response, whether or not that response looks like an error.
What a block response actually looks like
Anti-bot systems and WAFs have a few common ways of turning a request away without an error code. None of this is exotic or secret; it’s documented behavior from vendors like Cloudflare, Akamai, and PerimeterX, and it shows up consistently in scraping logs:
- A challenge page, often small and templated, with a fixed page size and recognizable strings (“checking your browser”, “verify you are human”, a specific CAPTCHA vendor’s script tag).
- A redirect to a verification or login wall that wasn’t part of the normal page flow.
- A response that’s technically the right page but missing the content you scrape for, because it was rendered server-side differently for a request that looked automated.
- Rate-limited soft throttling, where the server serves a cached or stripped page instead of an outright 429.
Each of these is detectable if you know what to look for on that specific target, and each is invisible to a monitor that only checks status codes.
Building a block signature per target
The starting point is target-specific, because block pages aren’t standardized. For a given domain, you build a small signature: a page-size range that’s abnormal for real content, a string or CSS selector that only appears on the block page, or a check that the fields you normally parse out are missing entirely. None of this requires guessing at how the target’s detection works internally. You’re reading the output you already get back, and classifying it.
A practical version looks like this: for every response, run your normal parser. If the parser can’t find the fields it expects, or the response matches one of your known block signatures, tag that request as blocked. This is a separate flag from HTTP-level success or failure. A request can be a 200, a parser failure, and a block, all at once, and that combination is exactly the case your error-rate alert was blind to.
Defining the metric
Block rate is simple once you have the tag: blocked requests divided by total requests, over a rolling window, broken out per target domain and ideally per proxy pool. The per-pool breakdown matters because a block rate spike is often localized. If block rate rises only on residential exit nodes from one region, or only through one proxy provider, that tells you something different than a rise across your entire fleet.
Alert thresholds should be set relative to that target’s own baseline, not a single global number. A target that’s aggressive about bot detection might run a background block rate that looks alarming compared to a quiet target, and that’s fine as long as it’s stable. What you’re watching for is a change: block rate that was sitting around a steady figure for weeks and then jumps. That jump is the signal, not the absolute number.
What a rising block rate should trigger
The point of this alert isn’t to hand you a way around the block. It’s to tell you, quickly, that something changed on the target’s side or in your own request pattern, so a human or an automated policy can respond sensibly. A responsible response to a rising block rate generally looks like backing off, not pushing harder:
- Reducing request concurrency and rate against that target until the block rate settles.
- Widening the interval between requests to a given endpoint rather than trying to squeeze more through.
- Pausing the job entirely and routing it for manual review if the block rate stays elevated after backing off, since that usually means the target changed its detection logic rather than just reacting to load.
- Checking whether the spike is isolated to one proxy pool, which more often points to that pool being flagged or exhausted rather than a change on the target’s end.
None of this is a guarantee that a scraper resumes working, and it shouldn’t be treated as one. A target is entitled to decide who accesses its site and how. The purpose of the alert is to keep your own operation honest about what’s actually happening, not to arm you for a fight you’re going to win by outlasting the other side.
Why proxy type changes your baseline, not your ceiling
Datacenter IPs, residential IPs, and mobile IPs all carry different reputational weight with anti-bot systems, because ASN and IP history are part of what these systems weigh. That’s well documented by proxy providers and anti-bot vendors alike. This means your background block rate on the same target can differ meaningfully depending on which pool you’re running through, and that’s worth baselining separately rather than averaging into one number. It does not mean any of these proxy types is undetectable or immune from being blocked; all three show up in block logs regularly, just at different rates depending on the target and how the pool is used.
If you’re comparing proxy providers, block rate broken out by target is one of the more honest metrics you can look at, because it’s built from your own logs against your own traffic rather than a vendor’s marketing claim. A provider whose pool holds a lower block rate on a target you actually scrape is more useful information than a general claim about IP quality.
Keep the error alert too
None of this replaces error-rate monitoring. Timeouts, connection resets, and outright 403/429 responses are still real signals worth paging on, especially because a sudden jump in hard errors can mean a proxy pool has gone down or a target has started blocking at the network level instead of serving a challenge page. Block rate sits alongside error rate, not in place of it. The two together give you a much more honest picture of whether a scraper is actually doing its job than either one alone.
Scraping cleanly at scale means treating the target’s response as data, not just a pass or fail. A block page returned with a 200 status is still telling you something, and it’s worth listening.
If you’re comparing proxy pools or building out monitoring for a scraping operation, you can find the rest of our guides and honest provider write-ups on the home page.
Get new guides and videos first — join the Telegram channel.