Pagination traps that quietly duplicate rows
The dataset looks fine until you count it
The first sign is usually a row count that doesn’t match expectations. You paginated through 40 pages at 50 rows each, so you expect 2,000 rows, but your database has 1,850 unique rows and 2,000 total inserts. Nothing crashed. No errors logged. The scraper ran clean from start to finish. The duplicates came from somewhere quieter: the pagination itself.
We’ve chased this exact discrepancy on production crawls more than once, and it’s rarely the scraper’s parsing logic. It’s almost always a mismatch between how the target site keeps track of “where you are” in a result set and how the scraper’s request pattern, especially proxy rotation and concurrency, interacts with that state. This is worth understanding in detail, because pagination duplicate results are one of the few data-quality bugs that won’t show up as an error and won’t get caught unless you’re specifically checking for it.
How pagination keeps state in the first place
There are three common approaches, and each fails differently.
Offset-based pagination (?page=3 or ?offset=100&limit=50) asks the backend to run the same underlying query and skip the first N rows. This only stays consistent if the underlying dataset and its sort order don’t change between your requests. If new rows get inserted at the top of a feed sorted by recency, or if rows get deleted between your page 1 and page 2 requests, the whole set shifts. A row that was at position 51 when you fetched page 1 can slide into position 50 by the time you fetch page 2, and you collect it twice. A row at position 100 can slide out of your window entirely and never get collected at all.
Cursor-based pagination (?after=abc123) asks the backend to resume from a specific row’s identity rather than a numeric position. This is more resistant to insertions and deletions because the cursor points at a row, not a slot. It’s still not immune, though: if the sort key used to generate the cursor isn’t unique (sorting by timestamp alone, for example, when multiple rows share a timestamp), rows on either side of a tie can end up duplicated or dropped depending on which one the cursor happened to land on.
Page-token pagination, common in API-driven sites, hands back an opaque token that encodes server-side state, sometimes tied to a specific backend node or a snapshot of the query at the time the token was issued. This is the most stable of the three when used correctly, but it’s also the most fragile if the token expires or if the session it was issued under gets invalidated mid-crawl.
Where rotating proxies introduce drift
This is the part specific to scraping at scale, and it’s worth being precise about the mechanism rather than hand-waving it.
Many sites bind pagination or session state to something beyond a cookie: the requesting IP, a fingerprint, or a combination checked at the load balancer. This isn’t always about anti-scraping measures either. It’s often just how session affinity works at the infrastructure level, routing a client to the same backend node or cache shard for consistency and performance.
If a scraper fetches page 1 through one proxy IP and page 2 through a different IP a few seconds later, because the proxy pool rotates on every request, the site may not recognize the second request as a continuation of the same session. Depending on how the backend is built, that can mean the “next page” request actually gets served against a fresh query, a different node’s cached snapshot of the data, or a re-sorted result set. Any of those can produce rows that overlap with what page 1 already returned. From the scraper’s point of view, both requests succeeded and both look like legitimate data. The duplication is invisible unless you’re diffing row identities.
This is a session consistency problem, not a detection-evasion problem, and it’s worth being clear about that distinction. A scraper that rotates IPs per request for load distribution across a pool is doing something completely ordinary. The failure mode described here is a side effect of that pattern colliding with how a specific site’s backend happens to be built, and it shows up whether or not the site has any bot-detection system running at all.
Concurrency makes the same problem worse
Even with a single stable IP, fetching pages concurrently instead of sequentially reintroduces the offset-drift problem. If you fire off requests for pages 1 through 10 at the same time against a feed that’s actively being written to, each request’s view of the “current” dataset can differ slightly from the others depending on exactly when it hit the backend. Two concurrent requests can both consider the same row to be at the boundary of their page, and both include it.
Sequential pagination against a live dataset isn’t perfectly safe either, but it at least narrows the window during which the underlying data can shift between requests.
Retries silently compound the problem
A block page, a CAPTCHA challenge, or a rate-limit response often comes back with an HTTP 200 status and a normal-looking page shell. If a scraper’s retry logic checks status code but not actual content, it can log “page fetched successfully,” move its cursor forward, and never come back for the real data on that page. Worse, if the retry logic instead treats a failed or ambiguous response as “try this page again,” and the underlying offset has since shifted, the retry can return an overlapping set instead of the original one. Either way, the fix is the same: validate that a response actually contains the expected data shape and row count before advancing pagination state, and treat challenge or block pages as a signal to pause and confirm state, not as a normal response to parse.
Building pagination that catches drift instead of hiding it
A few practices catch this class of bug before it pollutes a dataset:
Dedupe on identity, never on position. Use a stable field from the row itself (an ID, a URL, a hash of the content) as the uniqueness key downstream. Position in a result set is not an identity and should never be treated as one.
Keep proxy sessions sticky for the length of one pagination run. If a site’s backend is stateful across pages, using a single IP or session for the full run of one crawl job, then rotating between jobs, keeps the request pattern aligned with how the site expects to see it. This is a session-consistency choice, not a stealth technique, and it doesn’t guarantee the site’s backend is itself consistent.
Log expected versus actual row counts per page. If a page is supposed to return 50 rows and returns 37, that’s a signal worth surfacing immediately rather than silently appending whatever came back.
Prefer an explicit, unique sort key when the site’s API allows one. Sorting by a single non-unique field like a timestamp invites ties that cursor-based pagination can mishandle. Adding a secondary tiebreaker field, when available, removes the ambiguity.
Treat challenge and block responses as a distinct state, not as data. A response that doesn’t match the expected page structure should stop the pagination advance and get investigated, not get parsed as if it were a normal page.
When it’s the site, not the proxy setup
Worth saying plainly: if duplicates still show up with a fully sticky session and sequential requests, the cause is on the site’s side, usually a backend that isn’t internally consistent across its own replicas or that reorders results based on factors outside your control. No proxy configuration fixes that. What a well-built crawl can do is detect it quickly through row-count and identity checks, rather than letting it silently accumulate across a multi-day run.
Pagination duplicate results are a data engineering problem wearing a scraping costume. The fix isn’t a smarter proxy. It’s treating pagination state with the same care you’d give any other piece of distributed system state: verify it, don’t assume it holds still just because the HTTP status came back 200.
If you’re building out a proxy setup for a production crawl and want to think through session handling, rotation strategy, or how to match proxy type to the job, that’s exactly what we write about here.
Get new guides and videos first — join the Telegram channel.