← all guides

When a target starts serving stale cached pages

The symptom

You run a scrape, the request succeeds, the status code is a clean 200, and the page looks fine. Then you notice the price hasn’t changed in three days, or the “in stock” badge is showing for a product you know sold out this morning, or the timestamp on an article is from last week even though the site’s homepage shows something newer. Nothing errored. You just got old data back.

This is one of the more confusing things to debug in production scraping, because it doesn’t look like a block. There’s no captcha, no 403, no rate-limit header. The request worked. It just didn’t give you the truth. Understanding why this happens, and what it does and doesn’t mean about how a target is treating your traffic, matters more than most people think.

Caching exists between you and the origin, always

Almost no site you scrape at any real volume serves every request straight from its application server. Between the browser (or your scraper) and the database sits at least one caching layer, usually several: a CDN edge node, a reverse proxy like Varnish or nginx, sometimes an application-level cache, sometimes a full-page cache built into the CMS. These layers exist for a completely mundane reason: hitting the origin for every single request is expensive and slow, and most pages don’t change every second. So the infrastructure stores a copy and serves that copy to the next N requests until it expires or gets invalidated.

That’s normal, healthy web architecture. It has nothing to do with detecting you as a scraper. A cache doesn’t know or care who’s asking, at least not by default. It cares about the cache key: usually the URL, sometimes the URL plus a handful of headers the origin has told it to vary on (via the Vary header). If your request maps to a cache key that’s already populated, you get the stored copy. If the underlying content changed but the cache hasn’t been told to invalidate or hasn’t hit its TTL yet, you get a copy that’s technically correct at the protocol level and wrong at the business level.

Reading the headers instead of guessing

The honest way to know whether you’re looking at ordinary cache staleness or something else is to read what the response is telling you, because it usually tells you plainly.

Age is the header worth checking first. It reports, in seconds, how long the response has been sitting in a cache since it was fetched from the origin. An Age of 4 means you’re basically looking at a fresh copy. An Age of 40,000 means you’re looking at something over eleven hours old, and if that number keeps climbing across repeated requests instead of resetting, the cache in front of you isn’t refreshing on the schedule you’d expect.

Cache-Control tells you the policy the origin set, things like max-age=300 or stale-while-revalidate=60. If max-age is generous, thirty minutes or an hour, occasional staleness is just the site’s own tradeoff between freshness and load, not anything aimed at you.

ETag and Last-Modified let you do a conditional request: send If-None-Match or If-Modified-Since and the origin can reply with a 304 Not Modified if nothing changed, or a fresh 200 if it did. This is the cleanest way to distinguish “the content really hasn’t changed” from “I’m being served a stale copy of content that did change,” because a 304 is an honest, deliberate confirmation from the origin, not an inference you’re making from missing data.

X-Cache and similar vendor headers (CF-Cache-Status, X-Varnish, X-Served-By) often tell you outright whether you hit an edge cache (HIT) or the origin (MISS/BYPASS). These aren’t guaranteed to be present, plenty of production setups strip them before the response leaves the edge, but when they are there, they remove all the guesswork.

How proxy rotation interacts with this

This is where proxy-based scraping specifically intersects with caching, and it’s worth being precise about because it cuts both ways.

CDNs run many edge nodes distributed across regions and providers. Your request gets routed to whichever edge node is closest to, or otherwise assigned to, the IP you’re connecting from. If you’re rotating through a pool of proxy IPs that land you on different edge nodes each time, you can actually see more variance in freshness, not less, because each node has its own local cache with its own fill and expiry timing. One request through a Singapore-based residential IP might hit an edge that just refreshed. The next request, through a different IP that routes through a different edge, might hit one that’s sitting on a copy from twenty minutes ago. Neither is a fluke and neither is targeting you. It’s just the geography and cache topology of the CDN.

The flip side is also real: cache keys are sometimes coarser than people assume. If the origin’s Vary header doesn’t include anything IP-related (and it usually doesn’t, since caching by client IP defeats the purpose of caching), then everyone hitting the same URL through the same edge node gets the same cached object regardless of which proxy or which provider they’re using. Rotating IPs doesn’t change what’s in the cache. It only changes which edge node you’re asking.

When staleness is a signal, not an artifact

Ordinary caching explains most of what you’ll see. But it’s also true that some sites deliberately use a stale or simplified cached response as part of how they handle traffic they’ve classified as automated, serving a known-good, low-cost snapshot instead of hitting the live application for a request pattern that looks scripted. This is a legitimate operational choice on the target’s side: it protects backend capacity from a burst of non-human traffic without an outright block, which is often lower-friction for everyone than a hard 403 or a captcha wall.

From the outside, you can’t fully separate “this is just an edge cache doing its normal job” from “this is a deliberate stale response aimed at automated traffic” with certainty, and it isn’t the scraper’s place to try to force the origin to prove which one it is. Trying to defeat a caching layer, forging headers to fake a real browser, spamming cache-busting query parameters to force fresh origin hits, or hammering a target to force cache invalidation, is exactly the kind of behavior that turns a soft signal into a hard block, and it pushes load onto infrastructure that isn’t yours to spend.

What a compliant scraper actually does here

The reasonable response, once you notice a data source is consistently stale, is to slow down and read what’s available rather than escalate. Check whether the site publishes an API, an RSS feed, or a data export, since these are usually built specifically to serve current data without leaning on the cached HTML path at all, and using them is both more reliable and more respectful of the target’s infrastructure. Respect the Cache-Control and Age values you’re being given instead of polling faster than the stated max-age. Use conditional requests with ETag/If-None-Match so you’re only pulling a full response when something actually changed, which is lighter on the target and gives you a clean, origin-confirmed signal of freshness. And if a target’s terms of service prohibit scraping the page you’re looking at, that’s the boundary, not the cache header.

Where proxy choice matters in this picture isn’t about beating a cache, it’s about not making the situation worse. A pool of residential or mobile IPs spread across genuinely different networks will naturally land on different CDN edge nodes and produce natural-looking traffic patterns, which is simply lower-risk than a block of datacenter IPs firing rapid, identical requests at a single URL, a pattern that’s easy for both the CDN’s caching layer and any bot-detection sitting behind it to notice regardless of what data it decides to serve back.

Where this leaves you

Stale pages are usually a caching story, not a detection story, and the headers will tell you which one you’re looking at if you bother to read them. When they do point toward deliberate soft-blocking, the fix isn’t a workaround, it’s slowing down, checking for an official data source, and respecting whatever the target has published about acceptable use.

If you’re building out proxy infrastructure for scraping and want the honest tradeoffs between residential, mobile, and datacenter options, or an unvarnished look at where different providers actually hold up, that’s what we cover at Proxy Scraping.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →