← all guides

Reading robots.txt and the rules that actually bind you

Most people who scrape for a living have opinions about robots.txt they picked up secondhand. Someone told them it’s optional, someone else told them ignoring it gets you sued, and neither is quite right. The file itself is small and mechanical. What it means for you depends on what you’re doing with it.

What robots.txt actually is

robots.txt is a plain text file sitting at the root of a domain, like example.com/robots.txt. It follows the Robots Exclusion Protocol, a convention from the mid-1990s that predates almost every legal framework anyone tries to hang on top of it now. There’s no authentication, no cryptographic signature, nothing that forces a client to read it or obey it. It’s a note left on the door, not a lock on the door.

The file is organized into blocks, each starting with a User-agent line naming which crawler the rules apply to, followed by Allow and Disallow lines that say which paths that crawler can or can’t request. A wildcard user-agent (User-agent: *) applies to anyone not named more specifically elsewhere in the file. Some files also carry a Crawl-delay directive suggesting how many seconds to wait between requests, and a Sitemap line pointing at an XML sitemap, which is really there to help crawlers find content, not to police them.

Here’s a basic example:

User-agent: *
Disallow: /admin/
Disallow: /cart/
Allow: /products/

User-agent: Googlebot
Disallow: /search
Crawl-delay: 5

Sitemap: https://example.com/sitemap.xml

Read top to bottom, most specific match wins. If your scraper identifies itself with a user-agent string that matches Googlebot, the second block applies to you and the first one doesn’t. If it doesn’t match anything named explicitly, you fall under the wildcard block.

The technical binding: nothing stops you, and that’s the point

No server-side mechanism enforces robots.txt on its own. It is not a firewall rule. A scraper that ignores it entirely will not get an error for that reason alone. What robots.txt actually does is document intent, and that documented intent gets used in two very different ways depending on who’s reading it.

Well-behaved crawlers, meaning the ones written by companies that want to stay off blocklists and out of legal trouble, parse the file before crawling and route around disallowed paths automatically. Googlebot does this. Bingbot does this. Most commercial SEO and monitoring tools do this by default because their business depends on being seen as compliant.

The other reader is the site’s own anti-abuse tooling. Plenty of sites cross-reference request logs against their own robots.txt. A request that hits a disallowed path with a generic or spoofed user-agent is a stronger signal to a detection system than the same request against an allowed path, because it shows the requester either didn’t check the file or checked it and ignored it. That gets folded into whatever scoring the site uses to trigger rate limits, CAPTCHAs, or IP blocks. So robots.txt isn’t enforced by the protocol, but it’s frequently enforced downstream, by the same anti-bot systems you’re already trying not to trip.

This is where people get tangled up, because “is it legal to ignore robots.txt” and “is it technically enforced” are different questions with different answers.

robots.txt by itself is not a contract. You didn’t click anything, sign anything, or agree to anything by requesting the file. Courts that have looked at scraping cases (hiQ v. LinkedIn in the US being the most cited one, though its later procedural history softened the headline result) have generally treated robots.txt as a norm, not a binding legal instrument, when it stands alone.

What does carry legal weight, separately and often alongside robots.txt, is a site’s Terms of Service, particularly if you had to accept them to get access, and jurisdiction-specific computer access laws like the US Computer Fraud and Abuse Act. Scraping data behind a login wall after agreeing to ToS that forbid scraping is a different legal exposure than scraping public pages a robots.txt happens to disallow. The file itself doesn’t create the liability. It’s evidence of what the site operator communicated, and it gets used that way in disputes, which is meaningfully different from being the law.

None of this is legal advice, and if a scraping project touches personal data, paywalled content, or anything with real legal stakes, that’s a conversation for a lawyer who knows the relevant jurisdiction, not a blog post.

What actually binds you as an operator

Setting aside courtroom hypotheticals, here’s the practical stack that determines whether you should respect a given robots.txt rule:

The site’s own ToS. If it explicitly prohibits automated access or scraping, that’s the harder line, and it usually sits independent of robots.txt.

Rate and load signals encoded in the file itself. A Crawl-delay directive or a Disallow on a heavy, dynamically generated path (search results, filtered product listings, anything hitting a database on every hit) is frequently the site telling you, plainly, which parts of their infrastructure can’t take crawler load. Ignoring those isn’t just a compliance problem, it’s a good way to degrade the site for real users and get your whole IP range or ASN blocked as collateral damage.

Data sensitivity. robots.txt disallowing /user/ or /account/ paths is usually marking off personal data, not just server load. That’s worth respecting regardless of what the legal exposure looks like, because scraping personal data at scale carries its own separate risk under things like GDPR that has nothing to do with robots.txt.

What you intend to do with the data. Read-only competitive price monitoring on public product pages is a different risk profile than republishing scraped content wholesale or scraping behind authentication. The robots.txt file doesn’t distinguish between these, but you should.

A reasonable operating rule, and the one worth defaulting to: treat every Disallow as binding unless you have a specific, defensible reason not to, and never treat the absence of a rule as permission to hit a site harder than a careful human user would.

Reading the file for what it tells you about the site, not just the rules

There’s a second use for robots.txt that has nothing to do with compliance: it’s a map. A Disallow: /api/internal/ line tells you an API exists at that path. A block naming User-agent: Bingbot with different rules than the wildcard block tells you the site treats crawlers differently by identity, which usually means they’re checking user-agent strings somewhere in their stack, which in turn tells you spoofing a generic user-agent on a scraper is likely to get compared against that exact list. A Crawl-delay: 10 on a site that otherwise looks lightly protected is often the clearest signal you’ll get about their actual rate tolerance before other defenses kick in.

Reading it carefully before you write a single line of scraper code saves time later. It tells you where not to point requests, roughly how fast the site is willing to be crawled, and sometimes exactly which parts of the stack are sensitive enough that the site operator bothered to write a rule about them.

The proxy layer doesn’t change any of this

A rotating residential or mobile proxy pool changes what IP a request comes from. It does nothing to the legal or ethical status of the request itself. Routing a scraper that ignores Disallow: /account/ through a clean mobile IP doesn’t make that request more acceptable, it just makes it harder for the site to immediately block the source, which is a different thing entirely from making it right. Anyone building on Proxy Scraping content should already know we don’t sell proxies as a way to get past rules a site has clearly posted. They’re for scaling legitimate, permitted crawling without one flaky datacenter IP range taking down a whole job.

If you’re setting up scraping infrastructure and want more on the proxy side, rotation strategy, or how to pick between residential, mobile, and datacenter pools for different workloads, there’s more on the site.

Check out more breakdowns like this one on the Proxy Scraping home page.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →