Guides · Indexing

noindex vs robots.txt: crawling, indexing and the AI signals in between

8 min read · published 30 September 2026

robots.txt decides whether a crawler may fetch a page. noindex decides whether a search engine may show it. Mixing them up produces the two classic failures: private pages that appear in search results anyway, and noindex tags that nobody ever reads. A few newer directives try to say the same things to AI systems, with mixed success.

Crawling is not indexing

A search engine builds its index from pages it has fetched and from what other pages say about URLs it has not. If robots.txt disallows a URL, Googlebot does not fetch it, but it may still index the URL from links pointing at it, with a snippet like “No information is available for this page”. Google Search Console reports these as “Indexed, though blocked by robots.txt”. So robots.txt is the wrong tool for keeping something out of results.

noindex works the other way round. The crawler must be allowed to fetch the page in order to see the directive. If you disallow the page in robots.txt as well, the noindex is never read, and the page can stay indexed indefinitely. The rule of thumb: to remove a page from search, allow crawling and add noindex; to save crawl budget on pages you don’t care about either way, disallow them.

The two ways to say noindex

In the HTML head, for a page:

<meta name="robots" content="noindex, nofollow">
<meta name="googlebot" content="noindex">

As an HTTP header, for anything, including PDFs, images and API responses:

X-Robots-Tag: noindex
X-Robots-Tag: googlebot: noindex, nofollow

A crawler-specific tag or header applies only to that crawler; robots applies to all. The checker reads both and shows them under Signals, because a page can be perfectly crawlable and still invisible in search.

The AI-specific directives

SignalWhereSaysWho honours it
noai, noimageairobots meta or X-Robots-TagDo not use this page, or its images, for AIBegan at DeviantArt in 2022. Some hosting platforms emit it. No major crawler operator has documented support.
TDM-Reservation: 1HTTP header (also a tdm-reservation meta tag or a /.well-known/tdmrep.json file)Text-and-data-mining rights are reservedA W3C community-group spec (TDMRep) built for the EU copyright directive’s machine-readable opt-out. A legal statement rather than a technical block; it is what some European publishers and lawyers ask for.
Content-Signalrobots.txtWhether content may be used for search, AI input, AI trainingCloudflare’s 2025 proposal, emitted by its managed robots.txt. Explained in the Content-Signal guide.
Google-Extended, Applebot-Extendedrobots.txtWhether pages fetched for search may train or ground AIHonoured by Google and Apple, by their own documentation.

The pattern: the tokens the operators invented themselves are honoured; the ones invented on the publishers’ side are, at best, evidence of intent. That does not make them useless. Under the EU’s text-and-data-mining exception, an opt-out only counts if it is expressed in a machine-readable way, and TDM-Reservation exists precisely to be that expression.

Which one to use when

Run the checker after any change. The Signals cell shows noindex, nofollow, noai, TDM-Reservation and Content-Signal exactly as your server sends them, and the per-bot tiles show whether robots.txt and the server agree.

Check it on your site

More guides: How to block AI crawlers · robots.txt for AI crawlers · What is llms.txt · Content-Signal explained · Cloudflare blocking Googlebot.

Spotted something out of date? Operators change their bots often; email info@canaicrawl.com and we’ll re-check the guide.