How many requests at once is actually safe
A job that has to pull 100,000 pages in a day needs about 1.2 requests a second. If a page comes back in two seconds, that is three requests in flight.
Three.
I have lost count of the number of people running that job at 64 threads, because 64 was the value in the example they copied.
There is no safe concurrency figure that holds across targets, and anybody quoting you one is guessing. I sell mobile proxy lines for a living, so read the rest of this as a person with a financial interest telling you to buy fewer of them.
The number belongs to the target
Your safe rate depends on what the site runs in front of itself, what its own traffic looks like at that hour, and how close to capacity it already sits when you turn up.
That last one moves through the day. A target that ignores 30 parallel connections at 3am local can be at its limit by 9am, and the traffic shed first is whatever looks least like a paying customer.
Same scraper, same lines: I have held 90 in parallel on one site for six days without a complaint, and been pushed back to 2 on another inside an hour. Neither figure tells you anything about a third site.
Work out how little you need
The dial is the wrong end to start from. Start with the deadline.
Work divided by the time window gives a required rate. Multiply that rate by your average response time and you have how many requests need to be in flight to hit it. That is the entire calculation, and the answer comes out lower than people expect every time.
400,000 product pages, three days, four seconds a page through a headless browser. That is 1.5 pages a second, so six in flight. Not 200.
Then run deliberately under the answer. A job sitting at 40% of what it could do absorbs a bad hour without anyone noticing. A job pinned at its ceiling turns a bad hour into a missed deadline, and a missed deadline is what makes somebody reach for the concurrency setting at 2am.
Total concurrency is a meaningless figure
20 threads across 20 addresses and 20 threads on one address are different products, and a single concurrency setting hides which one you have got.
The site never sees your job. It sees each address separately, and it counts per address. Nothing in an HTTP request announces that 40 connections share an owner.
So the figure worth tracking is requests per address per second. Everything above that is bookkeeping.
One thread on each of 20 addresses reads like 20 ordinary visitors. 20 threads on one address reads like one visitor with 20 hands.
This is where people underbuy their supply. Somebody asks me for two lines and then describes a job wanting 60 requests in parallel. That works out at 30 a second on each address, which no home connection on earth produces.
The answer is more addresses at the same total rate. It costs more, I am the one selling them, weigh that however you like. It is still the answer.
A flat interval is a signature
One request every 2.0 seconds, forever, is not a pattern a person generates.
People read, stall, open a tab, wander off. The gaps between their requests scatter badly. A machine with a fixed sleep produces a standard deviation close to zero, and spotting that needs no fingerprinting and no clever model. It is arithmetic on timestamps.
Put the variance in on purpose.
A two second sleep with 100ms of jitter bolted on does not qualify. If your mean gap is two seconds, let the real gaps run from 0.5 to 5. Same throughput, different shape.
Variance buys you something else as well. Two workers on two addresses with identical pacing correlate perfectly, and correlated timing is how a set of unrelated looking addresses gets read as one operator.
Bursts fail where the average passes
200 requests spread evenly across a minute and 200 fired in four seconds have the same average.
Rate limiting does not run on averages. It runs on short windows, usually a few seconds, sometimes a minute. Your hourly figure is invisible to it.
So the job that sleeps 50 seconds and then discharges everything it owes gets refused, while the same volume dripped out survives the night.
Batching creates this without anybody choosing it. You collect a page of links, fire them all, wait, collect the next page. The output is lumps with gaps between them, and every lump is a burst.
The fix is a rate limited queue in front of the worker pool. Workers should be starved by the rate you set and never by how fast responses happen to arrive.
Walk up to the ceiling, then step back one
Start lower than feels sensible. One request in flight per address, wide random gaps, one hour of genuine work against the genuine target.
Record three numbers: median response time, 95th percentile, and the share of responses your parser actually pulled fields out of.
Double it. Another hour. Same three numbers. Keep going until one of them moves, because one of them always moves before anything gets refused.
Then drop back a rung and stay there. That is your number, for that target, this month.
A day of this replaces a fortnight of guessing, and it leaves you a baseline to compare against when the same job starts misbehaving in November.
Latency degrades before errors show up
Error rate is the signal everybody watches and the last one to arrive.
Latency moves first. A strained target queues before it refuses, because queueing is free and refusing is a decision somebody had to configure in advance. A median climbing from 400ms to 900ms while you went from 8 in flight to 16 is the ceiling telling you where it is, with nothing failed yet.
Watch the tail harder than the median. The 95th percentile moves earlier and further, and it decides how much of your worker pool is parked in a wait state at any given moment.
Back off at degradation. Waiting for the block means waiting until you are already in a penalty box, and getting out of one costs more time than the throughput you gained on the way in.
Blocks are unreliable signals anyway. Sites return thin pages, stale cache, soft notices carrying a 200. A monitor counting status codes will report a perfect night while your parser writes empty rows.
Never raise the dial because the target got slow
This is the reliable way to convert a working job into a blocked one, and almost everybody does it once.
The target slows. Maybe you caused it, maybe their database is having an afternoon. Throughput drops and the deadline stays where it was.
Look at what raising concurrency does there. A slow target holds every connection open longer, so at an unchanged thread count you already have more requests in flight than you did an hour ago. Raising the count multiplies a number that went up on its own.
And if your load contributed to the slowdown, you have just fed something that was already struggling. It gets slower, you add more, and the loop closes in about twenty minutes.
When a target slows, slow with it. Finish late.
What I got wrong
A job of mine pulled roughly 140,000 pages a night off one source. Eight mobile lines, five threads on each, clean for five weeks.
Then the window shrank from nine hours to six.
I moved to 12 threads a line, watched it for 20 minutes, saw no errors, and went to bed.
It had been failing since hour two. No errors involved. The responses looked healthy and my parser was finding about a third of the fields it used to, because the source had started serving a lighter version of the page to whatever it had decided I now was. Three nights of that went into the database before anybody spotted it, and cleaning it up took longer than the job.
Two mistakes, and the thread count was neither of them. I judged a change over 20 minutes when its effects took two hours to surface. And I was asserting on status codes instead of on extracted fields, so my monitoring said fine the whole way through.
The arithmetic was wrong as well. Nine hours down to six is a 50% increase in required rate. I went from 40 in flight to 96, which is 2.4x. The fix was five threads a line again plus four more lines, about $46 a month in SIM and modem depreciation, hitting the six hour window at the per address rate that had already worked for five weeks.
Nobody can hand you the number
I cannot tell you yours, and neither can any other supplier. A supplier who quotes one without knowing your target is quoting whichever figure sells the most lines.
The ceiling moves under you too. A site that tolerated 30 in parallel in March can tolerate 6 in August because they swapped a vendor, and measuring again is the only way you find out.
So measure on a schedule. Monthly is plenty for a stable source. Keep the numbers somewhere you can lay them side by side, because a slow drift stays invisible unless last month is written down next to this one.
The ladder I actually run, and the rest of what I use on my own lines, is here.
Get new guides and videos first — join the Telegram channel.