What changes when a target adds a CDN
I saved a failing response on a Tuesday and did not open it until Friday. The answer was four lines into the headers. A cache status field that had not been there a month earlier, and a server name belonging to nobody I was scraping.
Three days. Two extra proxy lines bought in the middle of it, on a theory that was already wrong when I paid.
The site itself had not changed. I diffed the html of a product page against a copy saved in March and every part my parser touched was identical. Same layout in a browser, same fonts, same cards. Nothing looked different because the difference was in front of the site.
Why I am the one telling you this
I run mobile proxy lines in Singapore. Sim cards, modems on a powered hub, roughly $10 a month per line in carrier data and another $1.50 in depreciation. So this failure tends to walk in with its diagnosis already attached, and the diagnosis is that the addresses went bad.
I sell the thing that diagnosis leads to. I still think it is wrong most of the time.
I have also put an edge layer in front of my own sites, because bandwidth costs money and caching is the cheapest way to stop paying for it. Which means I have used that console from the other side. It is a ranked table of who is asking for the most, a score beside each row, and switches underneath. Nobody writes code to act on it. You read a row and tick a box.
That asymmetry is the whole story. An afternoon of configuration on their side, and every scraper pointed at them starts from zero.
What actually changed at the wire
Before, dns handed your client an address belonging to their server. You opened a connection to that machine, their web server read your request and answered it, and the conversation had two participants.
After, dns hands back an address belonging to a network of machines that are not theirs. The nearest one answers. It terminates your connection, reads your request, decides what it thinks of you, and only sometimes bothers the origin at all.
Their server still exists and still runs the same code. You have stopped being the one talking to it.
Your handshake is read before your request exists
This is the part that catches people who did everything else right.
Opening an encrypted connection means sending a hello first, before any http exists. That hello lists protocol versions, cipher suites in a particular order, extensions in a particular order, curve preferences, and the protocols you would like to negotiate. Taken together it has a shape.
Browsers have shapes. A scripting library has one too, and it looks like nothing a browser has ever emitted.
Their original web server had no opinion about any of this. It was software for serving pages.
The edge is different software, written by different people, for a job that includes separating clients at the door. So your user agent claims one thing while your handshake says another, and catching that contradiction takes an equality check.
Your retry logic never ran. Your headers were never read. It was settled in the first packet.
How to tell this from a redesign
Two checks, both free.
The response headers change character. A server name you do not recognise. Cache directives that were absent last month. A request identifier. Sometimes a cookie pressed on you during a first visit you never asked for. The body can be byte for byte identical while the headers are a different document.
I keep the headers of one known good response per target in the repo beside the scraper now. Diffing that against today takes fifteen seconds. It is also the check I skipped for three days.
The second tell is uniformity. A redesign breaks one parser, because a person edited one template, so your listing crawl carries on while the product page selector goes quiet.
An edge layer breaks everything at once. Urls that share no template. The robots file. The 404 page. If you went from a 2% failure rate to 100% between two nightly runs, across pages with nothing in common, nobody redesigned anything.
And it follows the address rather than the page. That same url in your browser at home is fine.
Why nothing paged you
A challenge can come back carrying an ordinary status code with a challenge document in the body, so your success check passes on every request while the rows go to zero. I have made that argument at length in the piece on health checks and will not do it twice.
What matters here is narrower. When an edge layer shows up, the failure arrives in a shape your existing checks were never built to see, and the gap between it starting and you noticing gets measured in nights.
The same request from two exits genuinely differs
Geographic spread is the point of these networks. Your request lands on whichever machine is nearest to wherever you came from.
So a request from a Singapore line and an identical request from a line in Frankfurt reach two machines holding two cache states, sometimes running two rule sets, because rules can be scoped by region.
Cache age is the sneaky half. A page you pulled an hour ago through one exit can be older than the one you pull now through another, and that reads exactly like the site editing itself under you.
Debugging goes badly in those conditions. You change one thing, retest through whichever exit happened to be free, get a different result, and credit your change.
So pin the exit. One address, one variable. And when you compare a failing response against a stored good one, confirm both came from the same address or the comparison is telling you nothing at all.
The part that got better
Caching, and it is a genuine upside.
While a copy is fresh you are served by the edge machine and the origin is never contacted. It does not know you were there. Rules that exist to keep load off the origin have less reason to fire on a path that never reaches it, and in my experience cached pages are the most permissive surface on the whole property.
You can see which one you got, too. The response usually states hit or miss and how old the copy is, in the same headers you should be logging by now.
One of my jobs got faster after a target did this. Median response on the list pages went from somewhere around 900ms to about 120ms. All the pain was concentrated on the search endpoint, which nobody can cache.
So split the job by cacheability. Cached pages cost the target nothing and you can take them faster than before. Anything reaching through to the origin is the expensive traffic, which is exactly the traffic the new rules were bought to watch, so slow that right down.
What actually fixes it
Read the response first. Save one complete failing response, headers and body, and open it in a text editor. Your parser is the thing that stopped seeing, so stop asking it. Ten minutes of work that most people skip on the way to losing a week.
Then the handshake, and here is the honest part. Matching a real client’s handshake is the substantive fix, and it is much harder than changing a user agent. It is not a header. There is no string you can set.
Two routes count as honest work. Drive a real browser, so the handshake belongs to a browser because a browser produced it. Or use an http client whose encryption layer was built to present a browser’s hello instead of its own.
I am not naming vendors and I am not walking anybody through defeating a particular product’s checks. The mechanism transfers. The product names do not.
Then use what just got cheap. Sitemaps, feeds, json endpoints that render into a static page. A fair share of what people spin browsers up to collect is published deliberately in a file nobody reads.
And ask. An edge layer appearing usually means somebody just looked at a bandwidth bill. An email describing what you pull and how often, asking for an endpoint, lands better than it used to.
Sometimes the right answer is that you stop
A job running four workers out of a scripting library may now want a real browser per worker. Call it a few hundred megabytes of memory each, cpu that used to be free, and an eight minute run stretching towards an hour.
Price that against the data honestly. If the rows were worth $40 a month of proxy and a cron entry, they are probably not worth a box, a fleet of browsers, and somebody keeping all of it alive.
I have dropped two targets on that arithmetic. One of them I could likely get working with a week of attention, and I decided the maintenance after that week cost more than the data was worth. I still think that was correct, and I know people who would call it giving up.
What I cannot tell you
I cannot tell you what any particular target runs. I watch a handshake leave and a response come back and I infer from outside, and I have been wrong doing it.
And a handshake that matches a real browser today will not match in six weeks, because browsers ship every few weeks and the shape moves each time. Treating it as a fix is the error. It is either something you maintain permanently or something you decide is not worth maintaining.
The header diff I run, the cached versus uncached split, and the numbers I use to write a target off are here.
Get new guides and videos first — join the Telegram channel.