Guides · robots.txt

robots.txt for AI crawlers: the complete guide

11 min read · published 30 September 2026

robots.txt has been around since 1994 and became a standard, RFC 9309, in 2022. Most of what goes wrong with it comes from three rules people don’t know: named groups replace the catch-all, the longest match wins, and a server error means “block everything”. This guide covers the format, what each AI operator’s token does, and how to test the result.

Where it lives and how it is fetched

The file is /robots.txt at the root of a host, one per scheme and host. https://www.example.com/robots.txt does not apply to https://example.com/ or to a subdomain. It must be plain text, UTF-8, and crawlers only guarantee to read the first 500 KiB.

The status code matters as much as the content. A 2xx is parsed. A 4xx, including 404, means “no rules”: everything may be crawled. A 5xx or a timeout is the dangerous one. RFC 9309 says a crawler must treat an unavailable robots.txt as a complete disallow, and Google documents that it stops crawling the site while robots.txt keeps failing. A misconfigured redirect or a bot-protection challenge on /robots.txt can therefore take a whole site out of search.

The format

# A comment
User-agent: Googlebot
User-agent: Bingbot
Disallow: /admin/
Allow: /admin/public/

User-agent: GPTBot
Disallow: /

User-agent: *
Disallow: /tmp/

Sitemap: https://example.com/sitemap.xml

Which group applies

A crawler looks for the group whose user-agent token matches its own product token, compared case-insensitively. Googlebot/2.1 matches Googlebot. If a matching group exists, the crawler uses only that group. If none does, it uses the * group. If neither exists, everything is allowed. The groups are not merged, so this file blocks GPTBot from nothing except /private/, even though the catch-all blocks everything:

User-agent: *
Disallow: /

User-agent: GPTBot
Disallow: /private/

One documented exception: Applebot follows the Googlebot group when robots.txt does not name it. The checker models that.

Which rule applies

Within the chosen group, every rule whose path matches the URL is a candidate, and the one with the longest path wins. When an Allow and a Disallow tie on length, Allow wins. So Allow: /admin/public/ beats Disallow: /admin/ for anything under /admin/public/. Order in the file does not matter.

The AI tokens

Each operator publishes its tokens and what they govern. The bot directory keeps the exact user-agent strings and documentation links; in short:

OperatorTrainingSearch or answersOn request
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
GoogleGoogle-Extended (token)Googlebot–
AppleApplebot-Extended (token)Applebot–
Perplexity–PerplexityBotPerplexity-User
MetaMeta-ExternalAgent–Meta-ExternalFetcher
Common CrawlCCBot––
ByteDanceBytespider––

Two nuances matter. Google-Extended is a permission token, not a crawler: Googlebot still fetches the page for Search, and Google-Extended only says whether that copy may train Gemini or be used for grounding. Google documents that it does not affect Search or the AI Overviews shown in Search. And the on-request fetchers state in their documentation that they may bypass robots.txt because a person asked for the page, so a Disallow for them is advisory at best.

Mistakes that break it silently

A sensible default for 2026

For most sites: allow search engines and AI search, block training, keep the on-request fetchers, and decide about SEO tools on bandwidth. The generator writes exactly that with the “Block AI training” preset, and the checker shows the result per bot with the matching line number.

Check it on your site

More guides: How to block AI crawlers · What is llms.txt · noindex vs robots.txt · Content-Signal explained · Cloudflare blocking Googlebot.

Spotted something out of date? Operators change their bots often; email info@canaicrawl.com and we’ll re-check the guide.