Rotating proxies without losing your session
A ten minute sticky window under a job that takes twelve minutes. I have watched that bug twice, and both times it ate most of a day, because nothing about it errors.
The run completes. Every response carries a 200. Around minute ten the target stops recognising the scraper, and everything after that is a login page being parsed as data.
Rotation and session continuity pull in opposite directions. Almost every broken scraper I have been asked to look at had rotation turned up by somebody who never checked which requests needed to stay together.
Why I am the one telling you to rotate less
I sell mobile proxy lines. SIM cards on Singapore carriers, modems on a shelf in my flat, roughly $10 a month per line in carrier data and another $1.50 in modem depreciation.
Rotation is a feature I sell. So is holding an address still.
What a session is from the target’s side
Somebody arrives and the site needs to remember something about them across the next few requests. What is in the basket. Whether they proved who they are.
The protocol remembers nothing by itself, so the site issues a token. Usually a cookie, sometimes a header value. Call it a receipt.
Here is the part that gets skipped. When the site writes that receipt, it also records what it knew about whoever received it, and the address is the cheapest fact available. The receipt is a string plus a claim about where it was born.
The mismatch is the signal, and both addresses can be clean
Present that receipt from somewhere else and you have created a comparison that costs the target nothing to run.
A token issued to a line in one city, handed back three requests later from a line 600km away, then a third one, all inside 90 seconds. No fingerprinting involved. No scoring model. Two fields and an inequality.
The uncomfortable part is that neither address was suspicious. Both would pass any reputation check you care to name. What got flagged was the pair of them holding the same receipt, which means you can spend real money on better addresses and finish worse off than you started.
People do move address mid session and sites cope with it. You leave a cafe, your phone drops off the wifi onto the carrier network, and the basket is still there. What gets tolerated is one change, occasionally two, between networks with some plausible relationship.
One receipt appearing on nine networks in four minutes has no equivalent in the physical world. So the question was never whether an address can change during a session. It is how often, how fast, and whether the pattern has a shape a body could produce.
What must not rotate mid flight
Anything after a login. The moment you authenticate, the site has bound something to you, and every request from then on is a claim about that binding.
Any form with more than one step. Quote builders, address forms, anything where step two only makes sense because step one happened.
A cart. Items go in and the server holds them against a session it issued to a specific visitor.
Pagination carrying a cursor. This one catches people because from outside it reads as stateless. You see page 2 and page 3 in the URL and assume either is independently fetchable. Often that token is a pointer into a result set the server built once for one session.
The rule under all four: if the target did anything on the first request beyond handing you bytes, you are in a session, and married to the address you opened it on.
What can rotate as hard as you like
Independent fetches with nothing shared between them. A product page, a public listing, an availability poll.
The test takes ten seconds. Could I run these requests in any order, on any day, and get the same output? If yes, rotate as aggressively as your supplier permits. If not, that is a session and the tutorial you are reading does not apply.
Most real jobs are a mixture, which is where the damage comes from. The listing crawl is stateless, the detail pages behind the account are not, and one config gets written for both.
Bind the address to a worker, not to a request
The fix is structural and smaller than people expect. Stop treating rotation as something that happens to a request. Make it something that happens to a worker.
A worker picks up a job. It takes one exit address and holds it, with a fresh cookie jar and a fresh browser context if you are driving a browser. It runs that whole job on that address, then discards the jar, the context, the address and anything it cached, and takes the next job with a new set.
Rotation still happens, at whatever rate you were rotating at before. It just happens between jobs instead of inside one.
Nobody builds it this way because the request level version is one config line and the worker level version is an hour or two of code you write yourself.
One thing to know before you set it up. Your worker count is now also your session count, so that number is doing two jobs at once.
Rotate at boundaries: the end of an account, a search, a paginated set. Any point where stopping and never coming back would lose nothing. Inside a job there is no such point.
A rotation is a reset, and the half measure is the worst option
When you rotate, drop everything. Cookie jar, local storage, cached tokens, any browser state you were carrying. A new address is a new person, and a new person has never been to the site.
What people do instead is rotate the address and keep the jar, because that feels like preserving progress. It does the opposite: it drags the one artefact tying those requests together across the boundary that should have severed the link.
Three ways to run this, and they are nowhere near equivalent.
Hold one address and keep the jar, and you read as a person on a home connection using a website. Ordinary.
Rotate every request and destroy the jar each time, and you read as a crowd of unrelated first time visitors. Odd, since none of them has any history, but nothing connects them to each other.
Rotate the address and carry the jar, and you have handed the target one identifier that stitches every address you own into a single set. It is an inventory, signed with the same token.
That third one is the default you get when somebody switches rotation on over a scraper that already had a cookie jar. It is the most common configuration I see and the only one with no argument in its favour.
If you keep the jar, keep the address. If you drop the address, drop the jar.
Per request rotation is the wrong default
It is the default in almost every proxy tutorial and it should not be. It made sense when the only figure anybody counted was requests per address and the sites people scraped barely tracked sessions, which was about a decade ago.
Sticky should be the default. Hold an address for a unit of work, rotate between units, and enable per request rotation deliberately for jobs where you have verified there is no state. The window length is the figure to check when you buy: measure your longest single job, then buy longer.
I am aware how that sounds from a supplier. Per request rotation is what sells large pools, and large pools are what I sell. Four sticky lines invoice for less than a rotating pool of forty. I would still tell you to buy the four.
It never announces itself
When a session breaks, the site does not refuse you. It stops recognising you. You get a login page with a 200 on it, an empty basket, or page 1 of the results where page 4 should be.
Status code monitoring reports a clean night. Row counts look right, because a login page is still a row.
So assert on something that can only exist inside a live session. A username in the header, an order reference, a field that renders only once you are through the door. If the marker is missing, fail loudly even though the server said fine. I run that check on the first request after every rotation.
What I got wrong
A job of mine logged in, pulled 11 good pages, then downloaded roughly 2,000 copies of the same login form over five hours and reported success.
The rotation setting was half of it. The config was global, one proxy setup shared by every job in that project, so a change made for a stateless listing crawl landed on an authenticated job nobody was thinking about.
The other half was the success check. A status code and a row count, both of which a login page clears easily.
The afternoon of work: per job proxy config, worker level session binding, and one assertion on a field that only exists when logged in.
What still bothers me is the 11 good pages. It worked briefly, the smoke test passed, and it only went wrong at request 12 when the first rotation fired. Anything that runs for ten requests and then quietly stops is a session problem in my head now.
Where holding the address stops being enough
Some targets bind a session to more than the address. If the site fingerprints the browser and ties the token to that too, holding your address still is necessary and nowhere near sufficient. I will not pretend I have that one solved.
I also cannot tell you how any specific site does its binding. I infer it from outside, from what breaks when I change one thing at a time, and I have been wrong more than once. Anybody who tells you exactly what a target checks, without working there, is guessing.
The worker pattern written out properly, with the checks I run on the first request after a rotation, is here.
Get new guides and videos first — join the Telegram channel.