Blog

What Robots.txt Can and Cannot Enforce

The robots file is frequently described as controlling crawler access. It publishes preferences and has no enforcement mechanism whatever, and confusing the two produces predictable disappointments.

It is a request, not a control

A crawler fetches the file, parses it, and decides what to do. Nothing in the protocol prevents a client from ignoring it, and nothing in the server's behaviour changes based on it.

Compliance is a matter of the operator's policy. Major search and preview operators comply because their business depends on being welcome, and that incentive is the entire enforcement mechanism.

Clients with no such incentive simply do not fetch the file, and their requests are indistinguishable from any other request to the server.

Disallowed paths are published paths

The file is publicly readable by design, so listing a path in it announces that the path exists and that the operator would rather it were not visited.

For anyone looking for interesting areas of a site, the file is therefore a useful index rather than a deterrent, which is the opposite of what listing was intended to achieve.

Anything genuinely sensitive needs authentication. A path is protected when the server refuses unauthorised requests, not when a file asks politely that they not be made.

Crawl directives shape cooperating behaviour

Beyond exclusions, the file can express crawl delays and point to sitemaps, which cooperating crawlers use to pace themselves and to discover content efficiently.

This is where the file provides real value, because it lets a site manage the load that welcome crawlers place on it without blocking them.

Support for individual directives varies between operators, and unsupported directives are ignored silently, so behaviour has to be verified against logs rather than assumed from the file.

Exclusion is not the same as removal

A disallowed page will not be fetched by a compliant crawler, but it can still appear in an index if other pages link to it, since the exclusion prevents fetching rather than knowledge.

Preventing indexing requires a directive the crawler can only see by fetching the page, which means the page must be allowed in order for the instruction to be read.

Blocking a page while also trying to suppress its indexing is therefore self-defeating, and it is one of the most common configuration mistakes on the technical side of publishing.

Enforcement belongs at the server

Where access genuinely needs limiting, the mechanisms are authentication, rate limiting and verified-crawler checks applied by the server itself.

The robots file complements those by telling cooperative parties what is wanted, which reduces unnecessary load and unnecessary conflict. It is a coordination mechanism among willing participants, and it works well within that scope.