← all guides

Measuring proxy latency properly and why the average lies

Every proxy dashboard shows you a number. “Average latency: 340ms.” It looks precise. It looks useful. It is, on its own, close to useless for predicting how your scraper will behave in production.

I run proxy infrastructure and scrapers that sit on top of it, and the single most common mistake I see operators make is treating a mean response time as if it describes the pool. It doesn’t. A pool of proxies is not one thing behaving consistently. It is hundreds or thousands of individual exit paths, each with its own carrier, route, and load pattern, and the average smears all of that into a single misleading digit.

Why the average hides the problem

The mean is sensitive to nothing in particular and everything in general. If you have 100 requests where 95 come back in 150ms and 5 come back in 6000ms, your average lands around 435ms. That number tells you nothing true about either group. It doesn’t describe the fast majority, and it doesn’t describe the slow tail that’s actually going to time out your requests and stall your concurrency.

This matters more with proxies than with almost any other infrastructure component, because proxy latency isn’t a smooth curve. It’s bimodal or multimodal by nature. A residential exit on a strong home fiber connection might respond in 80ms. A residential exit routed through a congested mobile tower on the same pool might take 4 seconds, or hang until your client times out. Average those together across a rotating pool and you get a number that describes neither.

Datacenter proxies tend to have tighter, more consistent distributions because the underlying network path is stable and doesn’t depend on someone’s home router or a cell tower’s current load. Residential and mobile proxies have wider distributions because the exit node is a real consumer device or SIM sitting on infrastructure you don’t control and can’t smooth out. That’s the tradeoff you’re accepting when you choose IP type, and it shows up directly in your latency distribution, not just your average.

What percentiles actually tell you

If you want a number that predicts scraper behavior, look at p50, p90, p95, and p99, not the mean.

p50 (median) tells you what a typical request looks like. This is usually close to what people intuitively expect “average” to mean, and it’s a better single number than the mean because it isn’t dragged around by outliers.

p90 and p95 tell you what your slow requests look like. This is the number that determines whether your timeout setting is sane. If your p95 is 2.5 seconds and you’ve set a 2 second timeout, you are going to fail roughly one in twenty requests on latency alone, independent of any blocking or detection issue. That failure will look like a reliability problem in your logs when it’s actually a timeout configuration problem.

p99 tells you about your tail risk at scale. If you’re running 50,000 requests a day, your p99 slice is 500 requests. Those are the ones that eat your retry budget, hold open connections, and quietly inflate your run time even when everything else is healthy.

I look at p95 as the working number for setting timeouts and concurrency limits, because it’s stable enough not to bounce around from a handful of outliers, but honest enough to include the real slow tail instead of hiding it the way the mean does.

How to actually measure it

Don’t measure latency with a single request against a single target. That gives you a sample size of one against one endpoint’s current load, which tells you almost nothing about the pool.

A proper measurement setup looks like this:

Sample across the pool, not one IP. If you’re evaluating a rotating proxy pool, you need requests distributed across a meaningful number of distinct exit IPs, not repeated hits through whatever IP happens to be assigned first. Rotation behavior itself varies by provider, so confirm you’re actually getting new exits and not being pinned.

Sample across time. Network conditions, especially for mobile and residential, shift by time of day and by carrier load. A test run at 3am doesn’t tell you what your scraper will see at 2pm when the same cell towers are carrying peak consumer traffic. Run your measurement window long enough to catch that variation, ideally spanning at least a full day if the workload will run continuously.

Measure against your actual target, not a generic benchmark endpoint. A proxy’s round trip to a speed test server tells you about the proxy’s general connectivity. It doesn’t tell you about the latency to the specific site you’re scraping, which depends on peering, CDN edge location, and that target’s own response time. If you’re scraping a site that’s slow to respond regardless of your connection, no proxy will fix that, and your measurement needs to isolate which side the delay is coming from.

Separate connection time from response time. Total request time is made of TCP/TLS handshake time through the proxy, then the proxy’s connection out to the target, then the target’s actual processing and response time. If you only log total elapsed time, a slow target and a slow proxy look identical in your data. Most HTTP client libraries expose these phases separately if you ask for them; use that instead of a single wall clock number.

Log failures as failures, not as missing data points. If a request times out or the connection drops, that’s not a data point you exclude from your latency stats, it’s the most important data point you have. A pool that returns fast responses 90% of the time and drops the connection entirely 10% of the time will show a great average if you only average the successes, and that average will actively mislead you about reliability.

Why this matters for concurrency and timeout tuning

Once you have a real percentile distribution instead of a mean, two decisions get much easier.

Timeout values should sit above your p95, not above your p50. Setting a timeout at the median guarantees you’ll cut off a meaningful chunk of legitimate slow-but-successful requests, which then trigger retries, which then add more load to the pool, which then makes p95 worse. This is a feedback loop I’ve watched operators build by accident, tightening timeouts because “average latency looks fine” while their retry rate climbs.

Concurrency limits should account for tail latency, not average latency. If you size your worker pool assuming every request takes your average of 340ms, you’ll be surprised when a batch of p99 requests each hold a connection open for 6 seconds, backing up your queue far more than the math on the average suggested. Size for the tail you’ll actually see, not the center of the distribution.

A note on what latency does and doesn’t tell you about blocking

Latency and blocking are related but separate problems, and it’s worth keeping them apart when you’re diagnosing scraper issues. A slow response can mean network congestion on a residential or mobile exit. It can also mean a target is deliberately slow-rolling a response as part of its own defensive throttling, which is a legitimate and increasingly common technique sites use against automated traffic regardless of which proxy type sits behind the request. Treat unusually high latency from a specific target as a signal worth investigating, not as proof of anything on its own, and never assume a proxy provider’s marketing claims about speed or stealth substitute for your own measured percentiles against your own target.

None of this is about finding a trick to get around a site’s defenses. It’s about running honest instrumentation so you know what your infrastructure is actually doing, and building timeouts and concurrency settings around real numbers instead of a dashboard average that was never built to answer the question you’re asking it.

If you want more breakdowns like this on choosing and running proxy infrastructure honestly, head back to the homepage for the rest of our guides.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →