Guides · robots.txt
robots.txt for AI crawlers: the complete guide
robots.txt has been around since 1994 and became a standard, RFC 9309, in 2022. Most of what goes wrong with it comes from three rules people don’t know: named groups replace the catch-all, the longest match wins, and a server error means “block everything”. This guide covers the format, what each AI operator’s token does, and how to test the result.
Where it lives and how it is fetched
The file is /robots.txt at the root of a host, one per scheme and host. https://www.example.com/robots.txt does not apply to https://example.com/ or to a subdomain. It must be plain text, UTF-8, and crawlers only guarantee to read the first 500 KiB.
The status code matters as much as the content. A 2xx is parsed. A 4xx, including 404, means “no rules”: everything may be crawled. A 5xx or a timeout is the dangerous one. RFC 9309 says a crawler must treat an unavailable robots.txt as a complete disallow, and Google documents that it stops crawling the site while robots.txt keeps failing. A misconfigured redirect or a bot-protection challenge on /robots.txt can therefore take a whole site out of search.
The format
# A comment
User-agent: Googlebot
User-agent: Bingbot
Disallow: /admin/
Allow: /admin/public/
User-agent: GPTBot
Disallow: /
User-agent: *
Disallow: /tmp/
Sitemap: https://example.com/sitemap.xml
- A group is one or more
User-agentlines followed by rules. Consecutive user-agent lines share the rules that follow. DisallowandAllowtake a path prefix.Disallow:with nothing after it allows everything. Paths are case-sensitive.*matches any run of characters and$anchors the end, soDisallow: /*.pdf$blocks PDFs anywhere.Sitemapis a standalone line, not part of a group, and takes a full URL.Crawl-delayis honoured by Bing, Yandex and some others. Google ignores it.
Which group applies
A crawler looks for the group whose user-agent token matches its own product token, compared case-insensitively. Googlebot/2.1 matches Googlebot. If a matching group exists, the crawler uses only that group. If none does, it uses the * group. If neither exists, everything is allowed. The groups are not merged, so this file blocks GPTBot from nothing except /private/, even though the catch-all blocks everything:
User-agent: *
Disallow: /
User-agent: GPTBot
Disallow: /private/
One documented exception: Applebot follows the Googlebot group when robots.txt does not name it. The checker models that.
Which rule applies
Within the chosen group, every rule whose path matches the URL is a candidate, and the one with the longest path wins. When an Allow and a Disallow tie on length, Allow wins. So Allow: /admin/public/ beats Disallow: /admin/ for anything under /admin/public/. Order in the file does not matter.
The AI tokens
Each operator publishes its tokens and what they govern. The bot directory keeps the exact user-agent strings and documentation links; in short:
| Operator | Training | Search or answers | On request |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Google-Extended (token) | Googlebot | – | |
| Apple | Applebot-Extended (token) | Applebot | – |
| Perplexity | – | PerplexityBot | Perplexity-User |
| Meta | Meta-ExternalAgent | – | Meta-ExternalFetcher |
| Common Crawl | CCBot | – | – |
| ByteDance | Bytespider | – | – |
Two nuances matter. Google-Extended is a permission token, not a crawler: Googlebot still fetches the page for Search, and Google-Extended only says whether that copy may train Gemini or be used for grounding. Google documents that it does not affect Search or the AI Overviews shown in Search. And the on-request fetchers state in their documentation that they may bypass robots.txt because a person asked for the page, so a Disallow for them is advisory at best.
Mistakes that break it silently
- Serving HTML. Frameworks that answer every unknown path with the app shell return a 200 HTML page for
/robots.txt. Crawlers find no rules in it and treat the site as open. The checker flags this as “an HTML page, ignored”. - A UTF-8 byte-order mark or a stray character before the first
User-agentline, which can make the first group unparseable. - Rules before any group. An Allow or Disallow that appears before the first User-agent line belongs to no group and is ignored.
- Blocking assets. Disallowing
/static/or/_next/keeps Googlebot from rendering the page, which hurts indexing even though the HTML itself is allowed. - Expecting privacy. A disallowed URL can still appear in search results if other sites link to it. Only
noindexkeeps a page out of the index; see the noindex guide. - Trusting the file alone. A firewall can refuse the same bot robots.txt allows, and nothing in robots.txt shows it. That is why the checker sends a live request per bot.
A sensible default for 2026
For most sites: allow search engines and AI search, block training, keep the on-request fetchers, and decide about SEO tools on bandwidth. The generator writes exactly that with the “Block AI training” preset, and the checker shows the result per bot with the matching line number.
Check it on your site
More guides: How to block AI crawlers · What is llms.txt · noindex vs robots.txt · Content-Signal explained · Cloudflare blocking Googlebot.
Spotted something out of date? Operators change their bots often; email info@canaicrawl.com and we’ll re-check the guide.