Guides · Policy
Content-Signal explained: Cloudflare’s robots.txt policy for search, AI input and AI training
Since late 2025, many robots.txt files carry a line like Content-Signal: search=yes, ai-input=no, ai-train=no. It is Cloudflare’s Content Signals Policy: a way to say, in the file crawlers already read, what a site permits its content to be used for. It is a statement of terms, not a block, and it is worth understanding both halves of that sentence.
The syntax
Content-Signal is a line inside a robots.txt group, alongside Allow and Disallow, so it applies to the user agents that group names. Three signals are defined, each set to yes or no:
User-agent: *
Content-Signal: search=yes, ai-input=no, ai-train=no
Allow: /
- search: building a search index and showing results that link back to the source. Ordinary web search.
- ai-input: feeding content into an AI model at answer time. Retrieval, grounding and the summaries that AI search products show, whether or not they link back.
- ai-train: using content to train or fine-tune models.
A signal that is not mentioned expresses no preference. The order does not matter, names are case-insensitive, and a crawler that does not understand the line ignores it, because robots.txt parsers skip unknown fields. That is what makes it safe to add.
Where it comes from
Cloudflare published the policy in September 2025 as part of its push to let publishers set terms for AI companies, alongside its AI-crawler blocking and its pay-per-crawl experiment. Sites that use Cloudflare’s managed robots.txt get the line added automatically, at the time of writing expressing search=yes, ai-train=no and leaving ai-input unset. The policy text is deliberately written like a licence: it asserts that by crawling a site whose robots.txt carries the signals, an operator accepts the stated terms.
What it enforces
Nothing, by itself. Content-Signal is a preference expressed in a file. A crawler that respects it will not use your pages for the purposes you set to no; a crawler that does not respect it fetches exactly as before. No major AI operator has, as of this writing, publicly committed to honouring the signals, so its practical weight today is legal rather than technical: it is a clear, dated, machine-readable record of what you permitted, which matters if a dispute ever turns on whether an operator had notice.
That is the same position as the older TDM-Reservation header, which the noindex guide covers, and the two coexist happily. Content-Signal is more specific, distinguishing search from AI answers from training, which is the distinction most publishers actually want to draw.
How it interacts with your other rules
Content-Signal does not replace a Disallow. If you want training crawlers gone rather than merely told, keep the explicit groups for GPTBot, ClaudeBot, CCBot and the rest, and keep Google-Extended disallowed. A typical file that does both:
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
Disallow: /
User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Allow: /
Here search engines and AI search crawlers may fetch and use pages for search and answers, training crawlers are told to stay away and additionally denied permission to train, and anyone else reading the file learns the policy.
How the checker reports it
The checker parses Content-Signal lines per group and shows the first one under Signals as search=yes ai-input=no ai-train=no, next to noindex, TDM-Reservation and the sitemap. It does not change the score: the score measures what crawlers can fetch, and Content-Signal measures what you have told them they may do with it. Both are worth knowing; they are different questions.
Check it on your site
More guides: How to block AI crawlers · robots.txt for AI crawlers · What is llms.txt · noindex vs robots.txt · Cloudflare blocking Googlebot.
Spotted something out of date? Operators change their bots often; email info@canaicrawl.com and we’ll re-check the guide.