Guides · Indexing
noindex vs robots.txt: crawling, indexing and the AI signals in between
robots.txt decides whether a crawler may fetch a page. noindex decides whether a search engine may show it. Mixing them up produces the two classic failures: private pages that appear in search results anyway, and noindex tags that nobody ever reads. A few newer directives try to say the same things to AI systems, with mixed success.
Crawling is not indexing
A search engine builds its index from pages it has fetched and from what other pages say about URLs it has not. If robots.txt disallows a URL, Googlebot does not fetch it, but it may still index the URL from links pointing at it, with a snippet like “No information is available for this page”. Google Search Console reports these as “Indexed, though blocked by robots.txt”. So robots.txt is the wrong tool for keeping something out of results.
noindex works the other way round. The crawler must be allowed to fetch the page in order to see the directive. If you disallow the page in robots.txt as well, the noindex is never read, and the page can stay indexed indefinitely. The rule of thumb: to remove a page from search, allow crawling and add noindex; to save crawl budget on pages you don’t care about either way, disallow them.
The two ways to say noindex
In the HTML head, for a page:
<meta name="robots" content="noindex, nofollow">
<meta name="googlebot" content="noindex">
As an HTTP header, for anything, including PDFs, images and API responses:
X-Robots-Tag: noindex
X-Robots-Tag: googlebot: noindex, nofollow
A crawler-specific tag or header applies only to that crawler; robots applies to all. The checker reads both and shows them under Signals, because a page can be perfectly crawlable and still invisible in search.
The AI-specific directives
| Signal | Where | Says | Who honours it |
|---|---|---|---|
noai, noimageai | robots meta or X-Robots-Tag | Do not use this page, or its images, for AI | Began at DeviantArt in 2022. Some hosting platforms emit it. No major crawler operator has documented support. |
TDM-Reservation: 1 | HTTP header (also a tdm-reservation meta tag or a /.well-known/tdmrep.json file) | Text-and-data-mining rights are reserved | A W3C community-group spec (TDMRep) built for the EU copyright directive’s machine-readable opt-out. A legal statement rather than a technical block; it is what some European publishers and lawyers ask for. |
Content-Signal | robots.txt | Whether content may be used for search, AI input, AI training | Cloudflare’s 2025 proposal, emitted by its managed robots.txt. Explained in the Content-Signal guide. |
| Google-Extended, Applebot-Extended | robots.txt | Whether pages fetched for search may train or ground AI | Honoured by Google and Apple, by their own documentation. |
The pattern: the tokens the operators invented themselves are honoured; the ones invented on the publishers’ side are, at best, evidence of intent. That does not make them useless. Under the EU’s text-and-data-mining exception, an opt-out only counts if it is expressed in a machine-readable way, and TDM-Reservation exists precisely to be that expression.
Which one to use when
- Keep a page out of Google and Bing: allow crawling, add
noindex. Remove any robots.txt Disallow for it. - Keep crawlers off a section entirely (a staging area, search result pages, infinite calendars): robots.txt Disallow. Accept that URLs may still show up bare.
- Stay in search but out of AI training: Disallow the training crawlers and the Google-Extended and Applebot-Extended tokens in robots.txt. See the blocking guide.
- State a rights reservation: add
TDM-Reservation: 1as a header and Content-Signal lines to robots.txt. Keep the robots.txt blocks too; the statement does not enforce itself. - Keep a page out of everything: authentication. Nothing else is a lock.
Run the checker after any change. The Signals cell shows noindex, nofollow, noai, TDM-Reservation and Content-Signal exactly as your server sends them, and the per-bot tiles show whether robots.txt and the server agree.
Check it on your site
More guides: How to block AI crawlers · robots.txt for AI crawlers · What is llms.txt · Content-Signal explained · Cloudflare blocking Googlebot.
Spotted something out of date? Operators change their bots often; email info@canaicrawl.com and we’ll re-check the guide.