Bot directory
Every bot canaicrawl checks, who runs it, and the exact user agent we send. Each name links to a page on what the bot does and how to block or allow it. To build a full robots.txt, use the generator. To allow one bot your User-agent: * rules block, give it its own group:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Allow: /
This list is also published as JSON at /bots.json (CORS enabled, cached for an hour) with every bot’s token, user agent, group and documentation link, for scripts, firewalls and other tools. Please link back if you use it.
Seeing canaicrawl in your logs? Someone ran this checker on your site: one visit as canaicrawl/1.0 (and one as a browser if that was refused), one as each bot above, plus robots.txt, llms.txt and sitemap.xml. Questions: info@canaicrawl.com.
Search engines · Search index
The crawlers behind Google, Bing and the other search engines. Blocking one drops you out of its results.
| Bot | Operator | User agent we test with | Docs |
|---|---|---|---|
Googlebot |
Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html) |
Docs | |
Bingbot |
Microsoft | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/131.0.0.0 Safari/537.36 |
Docs |
DuckDuckBot |
DuckDuckGo | DuckDuckBot/1.1; (+http://duckduckgo.com/duckduckbot.html) |
Docs |
Applebot |
Apple | Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot) |
Docs |
PetalBot |
Huawei | Mozilla/5.0 (Linux; Android 7.0;) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; PetalBot;+https://webmaster.petalsearch.com/site/petalbot) |
Docs |
YandexBot |
Yandex | Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots) |
Docs |
Baiduspider |
Baidu | Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html) |
Docs |
AI search · AI answers
Crawlers that build the indexes AI assistants answer from and cite. Blocking one keeps you out of its answers.
| Bot | Operator | User agent we test with | Docs |
|---|---|---|---|
OAI-SearchBot |
OpenAI | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot |
Docs |
PerplexityBot |
Perplexity | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot) |
Docs |
Claude-SearchBot |
Anthropic | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-SearchBot/1.0; +https://www.anthropic.com) |
Docs |
DuckAssistBot |
DuckDuckGo | DuckAssistBot/1.0; (+http://duckduckgo.com/duckassistbot.html) |
Docs |
Amazonbot |
Amazon | Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/600.2.5 (KHTML, like Gecko) Version/8.0.2 Safari/600.2.5 (Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot) |
Docs |
YouBot |
You.com | Mozilla/5.0 (compatible; YouBot (+http://www.you.com)) |
Docs |
AI training · Model training
Crawlers that collect pages to train models. Blocking them does not affect search or AI answers.
| Bot | Operator | User agent we test with | Docs |
|---|---|---|---|
GPTBot |
OpenAI | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.1; +https://openai.com/gptbot |
Docs |
ClaudeBot |
Anthropic | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com) |
Docs |
Google-Extended |
robots.txt token only, never sent as a user agent | Docs | |
CCBot |
Common Crawl | CCBot/2.0 (https://commoncrawl.org/faq/) |
Docs |
Applebot-Extended |
Apple | robots.txt token only, never sent as a user agent | Docs |
GoogleOther |
GoogleOther |
Docs | |
Meta-ExternalAgent |
Meta | meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler) |
Docs |
Bytespider |
ByteDance | Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; spider-feedback@bytedance.com) |
none |
User-triggered fetchers · Live user request
Fetchers that load a page because a person asked an assistant about it. Blocking one means the assistant can’t open your pages on request.
| Bot | Operator | User agent we test with | Docs |
|---|---|---|---|
ChatGPT-User |
OpenAI | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot |
Docs |
Claude-User |
Anthropic | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-User/1.0; +Claude-User@anthropic.com) |
Docs |
Perplexity-User |
Perplexity | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user) |
Docs |
MistralAI-User |
Mistral | Mozilla/5.0 (compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots) |
Docs |
Meta-ExternalFetcher |
Meta | meta-externalfetcher/1.1 |
Docs |
SEO tools · Backlink & SEO index
Crawlers behind backlink and SEO tools. Many sites block them to save bandwidth; it has no effect on search.
| Bot | Operator | User agent we test with | Docs |
|---|---|---|---|
AhrefsBot |
Ahrefs | Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/) |
Docs |
SemrushBot |
Semrush | Mozilla/5.0 (compatible; SemrushBot/7~bl; +http://www.semrush.com/bot.html) |
Docs |
MJ12bot |
Majestic | Mozilla/5.0 (compatible; MJ12bot/v1.4.8; http://mj12bot.com/) |
Docs |
DotBot |
Moz | Mozilla/5.0 (compatible; DotBot/1.2; +https://opensiteexplorer.org/dotbot; help@moz.com) |
Docs |