Planning A Crawl From The Sitemap
Most scrapers start in the wrong place. They point a crawler at a homepage, follow every link they find, and end up with a queue that’s half navigation menus and half actual content. A sitemap gives you the site’s own map of what exists and roughly how important it is. If you’re running a scrape against any site with more than a few hundred pages, reading the sitemap first isn’t optional, it’s the difference between a crawl plan and a crawl that wanders.
I run proxy infrastructure for scraping jobs day to day, and the crawls that stay clean and finish on schedule are almost always the ones where someone sat down with the sitemap before writing a single request. This is about that planning step: what a sitemap actually contains, how to turn it into a queue, and how proxy pacing decisions should follow from that plan instead of being bolted on afterward.
What’s actually in a sitemap
A sitemap is an XML file, usually at /sitemap.xml, that lists URLs the site wants search engines to know about. Each <url> entry can carry a <loc> (the URL itself), <lastmod> (when it last changed), <changefreq> (how often it changes, though this field is widely ignored or wrong), and <priority> (a 0 to 1 relative weight the site assigns itself). None of these are guarantees. lastmod is the most useful and the most commonly stale, since a lot of CMS platforms stamp it at build time rather than actual content-change time.
Larger sites don’t ship one flat file. They ship a sitemap index, a file at the same location that lists other sitemap files, often split by content type: sitemap-products.xml, sitemap-blog.xml, sitemap-categories.xml. The spec caps a single sitemap file at 50,000 URLs or 50MB uncompressed, so any site past that size is required to split. Sitemaps are also frequently gzipped (sitemap.xml.gz), which your fetcher needs to handle before the XML parser ever sees it.
Before assuming /sitemap.xml is the only entry point, check robots.txt. Most sites declare their sitemap location there with a Sitemap: directive, and some sites use a non-default path or maintain more than one sitemap for different sections of the property. Reading robots.txt first also means you see the site’s declared crawl rules in the same pass, which matters for the pacing decisions later.
Turning the file list into a plan
Once you’ve pulled and parsed the sitemap (or the full tree of sitemaps, if there’s an index), you have a flat list of URLs. That’s not a crawl plan yet, it’s raw material. The planning step is deciding what order to hit them in and how to batch them.
A few practical groupings that actually help:
By URL pattern. Group URLs by path structure, /blog/, /product/, /category/, and treat each group as its own sub-crawl with its own parser. Sites change page templates by section, not globally, so a scraper tuned for product pages usually breaks silently on blog pages if you don’t separate them.
By lastmod recency. If you’re doing a repeat crawl of a site you’ve already scraped once, sort by lastmod and only queue what changed since your last pass. This is the single biggest efficiency gain available from a sitemap, and it’s the reason sitemaps exist for search engines in the first place: so crawlers don’t have to re-fetch everything to find what’s new.
By volume, against your own capacity. If the sitemap has 400,000 URLs and you’re planning to run this over a proxy pool sized for a few thousand requests an hour, you need to know that before you start, not on hour six when the queue is a quarter done and you’re deciding what to cut. Counting the sitemap upfront turns “how long will this take” into arithmetic instead of a guess.
Where the proxy pool decision actually fits
A sitemap-driven crawl tells you the shape of the job: how many URLs, how they’re grouped, how fast they change. That shape is what should drive your proxy setup, not the other way around.
If the crawl is a few thousand static pages read once, a modest datacenter proxy pool with sensible concurrency limits is usually enough, and it’s the cheapest option per request. If the target site serves different content or different markup to different regions, or if you need the crawl to reflect what a real visitor from a specific country sees, that’s a case for residential or mobile IPs, which route through real ISP or carrier networks rather than datacenter ranges. That’s a routing and geography decision, not a “residential beats datacenter” rule; the sitemap crawl itself doesn’t care what kind of IP fetched it, only that the response looked like what a normal page load returns.
Rotation policy follows the same logic. A crawl built around sitemap groupings naturally paces itself, because you’re moving through structured batches rather than firing every URL at once. That’s a better foundation for a rotating proxy setup than an unstructured link-follow crawl, where request timing is essentially random. If you’re spreading requests across a pool, matching your concurrency and rotation interval to the batch sizes from your sitemap plan keeps request timing closer to how a normal crawler or a human browsing session actually behaves, which is a byproduct of good planning, not a trick layered on top.
None of this makes a crawl invisible. Sites that care about scraping traffic look at more than IP reputation: request rate, header consistency, session behavior, and how closely a visit pattern matches a real browsing session all factor into how a site’s bot detection scores a client. No proxy type or rotation scheme removes that risk, and it’s worth being honest about that rather than pretending otherwise.
Pacing and politeness, not evasion
robots.txt often includes a Crawl-delay directive, and even when it doesn’t, treating the site’s own infrastructure with some care is just good practice, separate from any legal or contractual question of whether you should be scraping a given page at all. A sitemap-based plan makes this easy to implement, because you already know the total URL count and can set a target rate that finishes the job in a reasonable window without spiking request volume at any single moment.
Concretely: set a per-host concurrency cap, add jitter between requests instead of a fixed interval, and log response codes and response times as you go. A rising rate of 429s or 403s partway through a crawl is the site telling you, through its own infrastructure, that your current rate is a problem. The correct response is to slow down or stop, not to route around the signal. This isn’t a compliance formality, it’s also just better engineering: a crawl that gets itself blocked partway through has to be redone, which costs more time than pacing it correctly the first time.
Validating against what the sitemap actually claims
Sitemaps go stale. A page can sit in the sitemap for months after it’s deleted or moved, and lastmod can be wrong in either direction. Build a validation step into the plan: track 404 and redirect rates against what the sitemap claimed, and flag sections where the mismatch is high. A sitemap section with a 15% 404 rate is telling you the site’s sitemap generator is out of sync with its actual content, which changes how much you trust lastmod for prioritizing future crawls of that section.
It’s also worth checking sitemap freshness itself before trusting it for scope. Compare the sitemap’s own lastmod values against a spot check of a handful of pages you fetch directly. If the sitemap says a page changed last week but the page’s own content timestamp is six months old, that’s a sign the sitemap is templated rather than content-aware, and you should weight it accordingly in your planning rather than as ground truth.
A short checklist before you start
Pull robots.txt first and note any declared sitemap location and crawl-delay. Fetch the sitemap or sitemap index, handle gzip, and get an actual URL count before estimating runtime. Group URLs by path pattern and, on repeat crawls, by lastmod. Size your proxy pool and concurrency against the URL count and your target completion window, not the other way around. Log response codes as you go and treat rising error rates as a signal to slow down. Spot check a sample of pages against what the sitemap claims before trusting it for future runs.
None of this is complicated, but skipping it is the most common reason scrapes run long, get blocked partway through, or come back with a pile of stale and duplicate data. The sitemap is free information the site already published. Reading it properly before you write the crawler is the cheapest planning you’ll do on the whole project.
Get the full breakdown of sitemap parsing, proxy pool sizing, and rotation strategy on the Proxy Scraping home page.
Get new guides and videos first — join the Telegram channel.