Site redesign breaks scraper: how to keep production scraping jobs alive
A scraper that ran clean for eight months can stop working overnight without a single error in the logs. No blocked requests, no captchas, no rate limiting. Just empty fields where prices used to be, or a JSON payload that no longer matches your schema. Nine times out of ten, the cause isn’t detection. It’s a redesign. The site shipped new markup, and your selectors are now pointing at nothing.
This is one of the most common ways production scraping jobs fail, and it’s also one of the most avoidable if you build for it from the start instead of firefighting it after the fact.
Why a redesign breaks a scraper that was working fine
Scrapers depend on the structure of a page, not just its content. If you’re pulling a price out of div.product-price-value, you’re betting that class name stays put. Frontend teams don’t think about that bet at all. A CSS refactor, a move from server-rendered HTML to a client-side framework, a new component library, an A/B test that ships to 100% of traffic. Any of these can rename, restructure, or relocate the exact nodes your parser depends on, and none of them are aimed at scrapers.
Auto-generated class names make this worse. Modern build tools often hash CSS classes (class="a3f9k2") and regenerate those hashes on every deploy. If your selector locked onto a hashed class instead of a semantic attribute, it can break on the next routine deploy even without a visible redesign.
Structural changes are the second failure mode. Pagination that used to be page-numbered URLs becomes infinite scroll with a “load more” button that fires a background API call. A table becomes a grid of cards. Your XPath that counted “the 3rd td in the 2nd tr” now points at nothing, because there’s no table at all.
The signals that show up before the full break
Total failure is the easy case, because it’s loud. The harder case is partial breakage: the scraper still runs, still returns 200s, still writes rows, but half the fields are now empty or wrong. That kind of failure can sit undetected for weeks if nothing is watching the shape of the output, not just whether the job completed.
Watch for a rising null rate on fields that used to be reliably populated. Watch for output row counts that drop or spike without a matching change in the site’s actual inventory. Watch for values landing in the wrong type, like a price field suddenly containing a shipping estimate because the DOM order shifted. These are the tells that a redesign has partially landed, before it fully breaks your pipeline.
Writing selectors that bend instead of snap
The single highest-leverage habit is choosing what you key your selectors on. Presentation attributes (class names, nesting depth, inline styles) change with every redesign. Semantic attributes change far less often, because they’re often load-bearing for the site’s own functionality, not just its look.
In practice that means preferring:
data-*attributes,itemprop, and schema.org microdata over class names, when they exist- IDs that map to a stable backend concept (
id="product-42981") over positional selectors - Structured data blocks (JSON-LD in a
<script type="application/ld+json">tag) over parsing the visible DOM at all, when a site publishes them - Text-anchored matching (“find the label ‘Price’, then read the sibling node”) over pure positional matching, when structured data isn’t available
None of this makes a scraper immune to breakage. A redesign can remove JSON-LD entirely or restructure the microdata. But it meaningfully cuts how often you’re rewriting parsers, because you’re depending on things the site has less reason to churn.
Build fallback chains rather than a single selector path. If the primary selector returns nothing, try a secondary one before failing the field. Log which path fired. That log becomes your early warning system: a jump in secondary-path usage tells you the primary path is degrading before it fails completely.
Where proxies fit and where they don’t
It’s worth being precise about this, because it’s a common point of confusion. Proxies solve an IP-level problem: how many requests you can make from a single address before a site’s infrastructure starts rate-limiting or blocking that address, and how closely your traffic’s origin resembles a real user’s. Rotating through a pool of residential or mobile IPs, versus a block of datacenter IPs, changes how your request volume looks at the network and reputation layer.
A redesign is a completely different layer. It changes the HTML, not the network path. Swapping proxies, rotating IPs faster, or moving from datacenter to residential exit nodes does nothing to fix a broken selector, because the request is still succeeding and returning a 200. The page just doesn’t look the way your parser expects anymore. Conflating the two is how teams waste a day tuning proxy pools when the actual fix is a ten-line change to a parsing function.
Where proxies do matter here is indirectly: if you’re running frequent structural-diff checks against a site (see below), doing that from a single fixed IP at high frequency is itself a pattern that can draw scrutiny under normal rate-limiting logic, separate from anything to do with the redesign. That’s a reason to keep your monitoring traffic at a sane, low-frequency cadence and within the site’s published rate limits and terms, not a reason to lean on proxies as a workaround for anything else.
Monitoring that catches breakage before your downstream data does
The teams that handle redesigns well aren’t the ones with the cleverest selectors. They’re the ones who find out fast. A few concrete practices:
Schema validation on every run. Define the expected shape of a record (types, required fields, plausible ranges) and validate output against it before it lands in your database. A price field returning a string instead of a number, or a null where 99% of rows have a value, should fail loud, not get silently written.
Sample-based visual or structural diffing. Periodically snapshot the raw HTML of a handful of pages and diff the DOM structure against the last known-good snapshot. You don’t need this on every page on every run; a small rotating sample is usually enough to catch a redesign within a day of it shipping, rather than a week.
Alert on rate-of-change, not just failure. A job that goes from 2% null rate to 40% null rate overnight is a different problem than one that fails outright, and it needs a different alert. Track the metric, not just the exit code.
Keep a change log per target site. When you do fix a selector after a redesign, note what changed and when. Over a year, this turns into a useful record of how often a given site restructures, which tells you where to invest in more resilient parsing versus where a simple selector is fine.
When it breaks anyway
It will, eventually, no matter how defensively you write the parser. When it does, the fix is almost always the same sequence: pull a fresh copy of the page, diff it against your last known-good structure, identify what moved, and update the selector chain. Check whether the site now exposes structured data (JSON-LD, an API the frontend calls) that it didn’t before, since redesigns sometimes introduce a cleaner data source than what you were parsing previously. And check the site’s current robots.txt and terms before resuming, since a redesign is also a reasonable point to confirm you’re still scraping only what’s publicly accessible and permitted, especially if the new frontend exposes account-gated or paywalled content that the old one didn’t clearly separate.
Building a scraper for the long run means accepting that the target page is not a fixed contract. It’s someone else’s product, and it will change on their schedule, not yours. The goal isn’t to make a scraper unbreakable, because nothing running against a live site is. The goal is to make breakage visible fast, cheap to diagnose, and cheap to fix, so a redesign costs you an hour of selector updates instead of a week of quietly corrupted data.
If you’re comparing proxy types for the network side of a scraping stack, or want a plain look at how residential, mobile, and datacenter pools actually differ in practice, you can find more on that at Proxy Scraping.
Get new guides and videos first — join the Telegram channel.