What a proxy error is actually telling you
I get a ticket most weeks that opens with “your proxies are down”. There is usually a screenshot attached, and a good share of the time that screenshot shows a 403 with the target’s own error page rendered inside it, branding and cookies and all.
Nothing on my side was down. The site said no, the refusal travelled back through my gateway, and my gateway collected the blame.
That mix up is the expensive one, because there are two machines in the path and both of them answer in the same language.
The only question worth asking first
Which box produced this response.
Everything else follows. A failure made by the proxy and a failure made by the target share no cause and no fix, and they want opposite retry behaviour. In a stack trace they are the same three lines.
Most people never ask, because the proxy is the component they paid for and the one they can replace with a card. So the proxy gets swapped and the failure rate sits exactly where it was.
407 is the proxy, always
The one unambiguous code in this business.
407 is proxy authentication required. A website cannot send it. It has no standing to demand proxy credentials from you and no software on that side produces one. If you are holding a 407, the box that made it is the box you pay.
Causes cluster into a few:
- credentials wrong or stale
- your current address missing from an allowlist
- a subscription that expired overnight
- credentials sent on the wrong scheme, socks5 against http
Which of those it is matters less than the fact that none of them improve on attempt two.
403 is the target, nearly always
A 403 is a refusal, and refusing visitors is a website’s job.
You confirm it by what came attached. A genuine 403 from a site arrives carrying that site’s furniture. Their cookies. Their cache headers. A body with their layout in it, or their standard error object shaped like the rest of their API.
A gateway that wants to turn you down does not normally reach for 403. It has 407 for auth arguments and 502 for routing ones.
The exception is worth carrying around. Some providers answer 403 when you request a country outside your plan, or a hostname sitting on their internal blocklist. Those look different: a small body, almost always JSON, and not one cookie from the target anywhere in the response, because the request never travelled far enough to pick one up.
429 is either, which is why it costs people weeks
Same status code, two opposite meanings.
From the target it means you asked too fast, and the remedy is to slow down. From your provider it means you went past the concurrency your plan allows, and slowing down changes nothing, because the constraint is on sessions held open at once and pacing does not reduce that number.
Two tests separate them.
The furniture test again. A real 429 usually carries a Retry-After header, often a set of rate limit headers naming the window and what is left of your budget, and a body written by somebody who works there. A provider 429 is bare: short JSON, a phrase along the lines of session limit reached, frequently a header with the provider’s name on it.
Then latency. A 429 from the target has a round trip inside it, because it went somewhere and came home. A 429 from the gateway lands in twenty or thirty milliseconds, because it never left the building.
Before HTTP happens at all
Connection refused, and a timeout that fires while the connection is still being established, belong to the proxy or to the route reaching it. No target was ever involved.
A timeout that fires after the request went out is a different animal. The bytes left your machine. Either the far end is slow, or it is holding the socket open on purpose and feeding you a byte every few seconds to keep one of your workers parked.
Almost every HTTP client accepts a connect timeout and a read timeout as separate values. Almost nobody sets them separately. One number covers both, the log says “timeout”, and the log has told you nothing.
Five seconds to connect and thirty to read is a fine place to start. The exact figures matter far less than a timeout that now names which half of the journey died.
A 502 from a proxy is not the site being down
A 502 from your gateway says it could not reach the destination. DNS failed on its side, or the TCP connect to the target’s port failed, or the target dropped it part way through.
That is a statement about the leg between the proxy and the site. The site can be perfectly healthy and serving everybody else.
Breadth separates the two cases. A 502 on every target at once is the provider’s upstream and there is nothing on your side to fix. A 502 on one hostname while everything else runs clean is that hostname, or the route to it from that particular exit. The second happens more than people expect on mobile addresses, since a carrier’s path to a given host is not the path your office connection takes.
The log record that makes any of this possible
None of the above helps if your logger writes one string.
Put the two legs in separate fields:
{
"url": "...",
"exit": "sg-m1-08",
"proxy_status": 200,
"target_status": 429,
"connect_ms": 340,
"ttfb_ms": 2180,
"error": null
}
Six fields, and failures start grouping and sorting. Which exits are bad. Whether target 429s cluster by hour. Whether connect time drifted upward before anything visibly broke.
I ran a job that logged exceptions as strings for six weeks. When the failure rate climbed I could not answer whether the bad runs were concentrated on a handful of exits, because the exit was never written into the line. I rebuilt the logging, ran it two days, and the answer was three addresses out of forty.
Retries, where a misdiagnosis becomes money
A retry policy that cannot tell the two apart will retry the wrong one.
407: do not retry, fail loudly. No count of attempts has ever fixed a configuration problem, and a worker pool will happily produce tens of thousands of identical 407s in an hour.
Target 429: back off, widen the gap on each attempt, honour Retry-After if it arrived. Going straight back in extends the window they are counting inside.
Provider 429: per request backoff achieves nothing, because the gateway is queueing you. Cut your worker count instead.
403: stop retrying on the same session and the same exit. The same everything gets the same answer.
Connect failure: retry immediately somewhere else. The one case where a fast retry is correct, since a dead route yields no information and moving off it is free.
Read timeout: one retry, more patience, same exit.
The retry handlers built into HTTP libraries assume a single server on the other end, where three attempts on any error was a sensible default. Put a second machine in the middle and that same rule becomes an efficient way to repeat the wrong action at speed. Turn it off and write your own switch on the classification. Mine runs about twenty lines and has saved me more than any provider I ever moved to.
People will call that overengineering for a scraper. I have seen the invoices from the other approach.
Why HTTPS narrows the question for you
Against a plain http target the proxy sees every byte and can rewrite any of them. A 403 it invented is as convincing as one the site actually sent.
HTTPS takes that away. Your client asks the proxy for a CONNECT tunnel, and once that tunnel is up the proxy is shovelling encrypted bytes it can neither read nor alter. Any status code reaching you after the tunnel is established came from the target. There is no third possibility.
The proxy still gets to speak twice. On the CONNECT itself, where a non-200 means auth or routing, and by hanging up. Most clients surface both of those as a tunnel error rather than as a status code, which is exactly the distinction you want.
The exception is a gateway terminating TLS on your behalf, which you would see in the certificate chain, and which is worth knowing about before you push a login through anyone’s box.
Errors that were manufactured
Some providers return their own failures wearing the target’s clothes. On plain http that is easy for them: a 403 or a 429 with a believable HTML body and no marking of any kind.
There is a cheap way to take a sample. Request a hostname that cannot exist: a random subdomain of a domain you own, or anything ending in .invalid. Nothing out there can answer it, so whatever comes back was written by the gateway. Now you know what their fabrication looks like, and you can hold it against your real failures.
Timing does the same work from another direction. Anything returning faster than your normal round trip to that region never left.
I have no test for the harder version, and I want to be straight about that. A provider that copies a real error body byte for byte and pads the latency to match is indistinguishable from outside. I only ever caught one, by accident: two different exits returned pages identical down to a timestamp printed in the footer, and no site on earth serves the same second to two requests four minutes apart.
The classification table and the log fields I start from are here.
Get new guides and videos first — join the Telegram channel.