← all guides

What 403 vs 429 vs 503 really means for your proxies

Three codes, three different problems

Anyone running scrapers at scale eventually gets a wall of red status codes and treats them all the same way: rotate the proxy and try again. That works often enough to become a habit, and the habit is wrong. A 403, a 429, and a 503 come from different parts of a target’s stack, get triggered by different signals, and call for different fixes. Treating them identically either wastes good IPs or keeps burning ones that were never the problem in the first place.

This is written from the operator side, not the attacker side. The goal here is reading the signal correctly so a scraping operation stays within a site’s terms and doesn’t hammer infrastructure it has no business hammering. None of this is about defeating protection systems. It’s about understanding what they’re telling you so you can back off, slow down, or fix a config error instead of guessing.

403 forbidden: the request itself failed a check

A 403 usually means the request was evaluated and rejected on identity or fingerprint grounds. Something about the request looked wrong before the server even got to rate limits. Common triggers on the target side:

  • TLS fingerprint (JA3/JA4) doesn’t match a real browser
  • HTTP header order or a missing header (like Accept-Language) doesn’t match the declared user agent
  • The IP is on a known datacenter or proxy range that the site has decided to block outright
  • A WAF or bot-management vendor (Cloudflare, Akamai, PerimeterX, DataDome) scored the session as automated

The key thing about 403 is that it’s often not about volume at all. You can get a 403 on your very first request from a brand new IP if the fingerprint is bad. Rotating to another IP from the same subnet or the same proxy type will usually just get you another 403, because the problem is the request signature, not the address.

What actually helps: matching your HTTP client’s fingerprint to a real browser (consistent header sets, correct TLS stack), and using an IP class the site doesn’t blanket-reject. This is exactly why residential and mobile IPs get sold at a premium over datacenter IPs. It’s not that they’re invisible. It’s that they sit in ranges that carry real subscriber traffic, so a site can’t just block the ASN without blocking its own users.

429 too many requests: the server is asking you to slow down

A 429 is a rate-limit response, and it’s the most polite of the three because it’s the server explicitly telling you it noticed the pace and wants you to change it. This is measured per some key, and that key is not always the IP:

  • Per-IP request rate
  • Per-session or per-cookie rate
  • Per-account rate (if you’re authenticated)
  • Fingerprint-based rate, where the site is tracking a device signature across IPs

If the limit is per-IP, rotating proxies genuinely resolves it, because the counter resets for a fresh address. If the limit is tracked by fingerprint or session cookie, rotating the IP does nothing and you’ll get 429s on the new IP just as fast, because the server is following a signal that survived the rotation. This is a common trap: teams see 429s, assume it’s IP-based, buy more IPs, and the error rate doesn’t move because the actual constraint was concurrency or session identity.

The correct read on a 429 is to slow your request rate and check whether the response includes a Retry-After header. Respecting that header is the compliant path and it’s usually the fastest way back to normal service, faster than rotating and hoping.

503 service unavailable: it might not be about you at all

A 503 is the odd one out because it doesn’t necessarily mean the target is defending against your scraper. It means the origin server, or something in front of it, couldn’t serve the request right now. Causes include:

  • The backend is overloaded from legitimate traffic and shedding load
  • A CDN or WAF is in a “challenge” or “under attack” mode and issuing 503s while it verifies traffic
  • Scheduled maintenance or a deploy is in progress
  • The site’s bot-management layer is holding the connection in a JS-challenge state and returning 503 until a check passes

Because a 503 can be entirely unrelated to your scraper’s behavior, the instinct to immediately rotate and retry can make things worse if the site is already struggling under load. The considerate response is exponential backoff: wait, retry with increasing delay, and cap the number of retries. If a 503 stays constant across every IP in a large pool, that’s a strong signal the origin is down for everyone, not blocking you specifically, and continuing to hit it faster is just adding to the load.

Why the same error can mean opposite things for different targets

None of these codes are standardized in practice the way the HTTP spec suggests. A lot of sites return 403 for what is functionally a rate limit, because they don’t want to leak information about their rate-limiting logic. Others return 429 for what is actually a hard IP ban, just to avoid tipping off the scraper that the block is permanent. Cloudflare’s own challenge pages have historically returned 403 for JS-challenge failures that have nothing to do with the IP’s reputation and everything to do with missing browser execution.

This means the status code is a starting hypothesis, not a verdict. The way to confirm it is to look at what changes the outcome:

  • If switching IP class (datacenter to residential, or residential to mobile) fixes it, the block was IP-reputation based, consistent with a 403 pattern.
  • If slowing the request rate fixes it, it was a rate limit, consistent with 429, even if the site returned a 403 or 503 status.
  • If waiting fixes it regardless of IP or rate, it was likely a transient server condition, consistent with 503.

Logging the response code alongside the IP, the request rate at the time, and whether a retry with no other change succeeded, turns this from guesswork into a diagnosis. Most scraping frameworks make this trivial to add as middleware, and it pays for itself the first time someone asks why the error rate spiked.

What this means for proxy pool design

If your scraping targets return a lot of 403s, more IPs in the pool doesn’t fix it if they’re the same class of IP the site already blocks. Residential or mobile IPs, sourced from a network with real subscriber diversity, address the fingerprint and reputation problem that 403 usually represents.

If your targets return a lot of 429s, the fix is architectural: request pacing, concurrency limits per IP, and respecting Retry-After. Buying more IPs to spread the same total request volume across a bigger pool can work, but only if the rate limit is genuinely per-IP. Test that assumption before spending on it.

If your targets return a lot of 503s, look at your own retry logic before touching the proxy pool at all. Aggressive immediate retries across a big proxy pool during a 503 storm is how a scraper turns “site is having a bad day” into “site’s ops team notices a spike in a specific traffic pattern,” which is not a good place to be even by accident.

None of this guarantees a scraper avoids blocks, and no proxy type or pool size makes a scraping operation undetectable. Bot-management vendors combine dozens of signals beyond IP and rate, and a well-run detection system can flag a session on behavior alone regardless of how clean the IP looks. What correctly reading these three codes gets you is a scraper that fails for the right reason, backs off when it should, and doesn’t waste a proxy budget solving a problem that was never about the IP.

If you want more breakdowns like this on how proxy-based scraping actually behaves in production, come see the rest of what we’ve written at Proxy Scraping.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →