Building a proxy health check worth running
Forty addresses checked once a minute is 57,600 requests a day aimed at one site. The scraper those addresses existed for was pulling around 30,000. My monitor was the bigger customer by a wide margin.
It was also wrong.
It fetched an echo service through each address and looked for a 200. Everything read green for six days while the job’s nightly row count fell by about two thirds. Both facts held at the same time. Every address could reach an echo service, and the site I actually scraped had stopped serving them anything.
I sell mobile proxy lines, so apply the usual discount to everything below. SIM cards on Singapore carriers, modems on a shelf, roughly $10 a month per line in carrier data and another $1.50 in depreciation. The broken monitor was mine, watching my own lines, and I believed it for six days.
The loop, not the trial
There is a separate exercise where you evaluate a supplier before paying. Registration lookups, reputation lists, latency from the machine that will run the job, a slice of real traffic against the real target. That is an evening of work and you do it once, before money moves.
This is the other thing. The loop that keeps running for as long as you own the pool.
Its question is much smaller. Which address should the next request avoid, right now.
A 200 is a statement about the transport
A status code tells you a request left your box and something on the far side answered it. Worth knowing. Also the entire extent of what it says.
Block pages return 200. Interstitial challenges return 200. The soft block, where a site quietly hands you a version of itself with the data stripped out, returns 200 with a content length that looks unremarkable if you are not comparing it to a real one.
The status line describes the pipe. Your parser reads the body. Nothing in HTTP forces those two to agree.
So an endpoint that reads your own address back to you is a connectivity test. Naming it a health check is why week three arrives as a surprise.
Probe the target, or something shaped like it
Send the check at the site you scrape.
Pick a page, and pick a dull one. A category listing, a help article, anything that has held the same URL for two years and will not move because a product went out of stock.
Not the home page. Home pages get personalised, get served as variants, get a promo banner dropped over them on a Tuesday. Your check will flap for reasons that have nothing to do with your proxy, and you will start ignoring it.
If hitting the real target is too expensive, or you would rather monitoring did not become a visible share of your traffic there, find a structural stand in. Another site behind the same protection vendor, with the same challenge behaviour. Weaker signal, and a different universe from an echo service.
Four things worth asserting on
A string that only appears on a genuinely served page. Choose it carefully. The page title survives on plenty of block pages, because the same edge serves both, and so do the cookie banner and the footer. Go down into the part of the document that carries data: a column heading, a line of legal text that only renders under the product table, something a refusal page has no reason to contain.
A size floor. Blocks are small and challenges are smaller. Take the median byte count of a few hundred known good responses, set the floor around 60% of it, fail anything below. That catches the whole family of failures where the page arrives correctly shaped and empty.
The URL you finished on. A client following redirects silently will report a cheerful 200 from a consent wall on another host entirely. Compare where you ended against where you aimed.
The address the target thinks it is serving. Sites leak this constantly, in a currency guess or a country string sitting in the markup. Set it against the address you believe you are leaving from. Disagreement means a leak, and a leak makes every other number in your pool wrong.
The assertions that will burn you
Do not hash the page. Do not pin an exact byte count. Do not tie the check to the specific fields your scraper extracts, however natural that feels, because the first redesign turns the whole pool red inside a minute and you lose a morning power cycling hardware that was fine.
Which is the rule underneath all of this. When every address fails at once, the check is the suspect.
Forty lines across three carriers do not break simultaneously. Something they all share changed, and the thing they share most is your code.
Health is a rate, not a reading
A single failed check carries almost no information. Sites drop connections. A carrier gives you a two second hole for a reason you will never learn. Quarantine on one failure and you will spend your life quarantining working lines.
Keep every result per address and read the last 20. Success rate across that window is the figure that matters.
What matters inside it is movement. An address that held 98% for three weeks and now reads 85% is on the way out, even though 85% clears any static threshold you would have set. An address that has been at 85% since the day you got it is a slightly worse address and it is fine.
So baseline each address against its own first week. A single threshold for the pool assumes the pool is uniform, and I have never bought a uniform pool.
Your monitor is a traffic source
Back to the 57,600, because volume is the smaller half of that problem.
The shape is the real one. Exactly on the interval, exactly the same path, from every address you own, indefinitely. If you set out to design the most legible pattern a bot could emit, you would land somewhere very close to a monitoring loop.
So check rarely and jitter it. Mine runs every 15 minutes with a few minutes of random spread, rotating across four dull URLs so no single path collects a metronome from the entire pool.
The cheapest check is the work itself
If a worker just fetched a real page through address 12 and the parser pulled rows out of it, that address is healthy. Better evidence than any probe you could write, because it is the exact request, on the exact page, at the concurrency you actually run at.
So write the outcome of real requests into the same table the health check writes into. One schema: address, timestamp, pass or fail, bytes.
That leaves the synthetic probe one job. Cover the addresses with no recent evidence. Idle for two hours, probe it. Served 40 good pages in the last ten minutes, leave it alone.
Doing it that way removed most of my monitoring traffic and improved the signal at the same time, which is not a trade that comes up often.
Quarantine, then retest on a back off
Failure takes an address out of rotation. It does not take it out of the pool.
Addresses recover. A refused mobile line rotates through the carrier’s pool and returns on a completely different address twenty minutes later. A residential exit reconnects. Even a burnt datacenter address falls off a list eventually.
Delete on failure and you are erasing your own capacity in slow motion. Nobody notices until the night the job is running on 11 lines out of 40 and taking until morning.
My rule: three consecutive failures pulls it. Retest at 5 minutes, 20 minutes, an hour, then every four hours. Two consecutive passes puts it back. Anything still quarantined after a day has stopped being a monitoring matter and become a ticket to whoever sold it to me, or a walk over to the shelf.
Alert on the pool, not the address
Per address alerts train you to ignore alerts. Three of forty red is a normal Tuesday.
Alert on the aggregate. What share of the pool is live, and how fast that share is moving. Mine pages me below two thirds, because under that the job is running at a concurrency nobody chose.
Thirty of forty going red inside ten minutes is one of two things. The target changed, or your egress changed. Both need a person out of bed, and almost nothing else on that dashboard does.
The echo test, demoted to stage two
I have been rude about echo services. I still run one.
When a real check fails there is exactly one question worth answering first. Plumbing or reception. Did the packets fail to arrive, or arrive and get turned away.
One echo request settles it. Echo fine and target failing means the line works and the site does not want you. Both failing means go and look at the hardware.
So keep it, fire it only after a real check has already failed, and stop calling it monitoring.
The three weeks I passed on block pages
The echo version was the first mistake. The second one was worse.
I moved the check onto the real target and asserted on a string, then picked the string off the cookie banner because it sat near the top of the markup and was easy to find.
The banner is served by the edge. So is the block page. The block page carried the banner.
For about three weeks that check passed on refusals, and it did more damage than the echo version ever had, because I believed it harder.
Now I pull the assertion from as deep in the document as I can reach, and I verify it by deliberately getting myself blocked once and reading what comes back. Twenty minutes of work, and the highest value twenty minutes in the whole build.
What I still cannot tell you
A stand in target tells you the stand in works. I cannot promise the correlation holds, and I have watched it break. Mine stayed green through a week when the real target had tightened on exactly the address range I was using.
The thresholds are mine and not yours. 60% of median size, three consecutive failures, a window of 20, a 15 minute interval. They came out of my targets over a couple of years. Yours will differ, and only a fortnight of your own data will tell you by how much.
A health check also never tells you why an address stopped working. It tells you which one to stop using, and when. The reason lives somewhere else, and chasing it inside the monitor is how a monitor turns into a second scraper.
The schema I use, the back off timings, and the way I choose an assertion string that survives a redesign are here.
Get new guides and videos first — join the Telegram channel.