# canaicrawl > A free checker that shows whether AI crawlers, search engines and SEO bots can read a website. It checks the site's robots.txt rules, sends a live request as each of 30 bots, reads indexing signals (noindex, noai, TDM-Reservation, Content-Signal, sitemap) and looks for an llms.txt file. A report for any site can be linked as https://canaicrawl.com/?url=example.com and a score badge for any site is at https://canaicrawl.com/badge/example.com.svg. Live requests are sent from canaicrawl's server with each bot's user agent, not from the bot operator's IP addresses. Contact: info@canaicrawl.com. ## Pages - [Checker](https://canaicrawl.com/): enter a URL to get a 0–100 access score for a chosen goal and a per-bot result with the matching robots.txt rule and HTTP response - [Bot directory](https://canaicrawl.com/bots): the 30 bots checked (Googlebot, Bingbot, GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot and more), their operators and exact user-agent strings - [robots.txt generator](https://canaicrawl.com/robots-txt-generator): choose which bots to allow or block and copy the resulting robots.txt - [llms.txt generator](https://canaicrawl.com/llms-txt-generator): write an llms.txt in the llmstxt.org format and check an existing one - [Guides](https://canaicrawl.com/guides): dated guides on blocking AI crawlers, robots.txt, llms.txt, noindex, Content-Signal and CDN blocking - [About](https://canaicrawl.com/about): how a check works, what it sends to a site, and the score's weights - [bots.json](https://canaicrawl.com/bots.json): the bot list as JSON (name, robots.txt token, user agent, group, documentation) - [Privacy policy](https://canaicrawl.com/privacy): what the checker and the monitoring sign-up do with data - [Terms of use](https://canaicrawl.com/terms): the terms the free checker is offered under ## Guides - [How to block AI crawlers](https://canaicrawl.com/guides/how-to-block-ai-crawlers): A practical guide to blocking GPTBot, ClaudeBot, CCBot, Bytespider and other AI crawlers with robots.txt and firewall rules, which bots to keep, and how to verify the block works. - [robots.txt for AI crawlers](https://canaicrawl.com/guides/robots-txt-for-ai-crawlers): How robots.txt really works under RFC 9309: groups, precedence, wildcards, status codes, the AI bot tokens for OpenAI, Anthropic, Google, Apple and others, and the mistakes that silently break it. - [What is llms.txt](https://canaicrawl.com/guides/what-is-llms-txt): What the llms.txt proposal is, what the format looks like, which AI systems actually read it, and an honest view of whether adding one changes anything. - [noindex vs robots.txt](https://canaicrawl.com/guides/noindex-vs-robots-txt): The difference between blocking crawling and blocking indexing, when to use robots.txt, noindex and X-Robots-Tag, and what the noai, TDM-Reservation and Content-Signal directives do. - [Content-Signal explained](https://canaicrawl.com/guides/content-signal-explained): What the Content-Signal lines in robots.txt mean, where they come from, how search=yes, ai-input=no and ai-train=no are interpreted, and what they do and do not enforce. - [Cloudflare blocking Googlebot](https://canaicrawl.com/guides/cloudflare-blocking-googlebot): How Bot Fight Mode, managed challenges, WAF rules and the AI-bot toggle end up refusing Googlebot and other crawlers, how to see it in Cloudflare and Search Console, and how to fix it without turning protection off. ## Bot pages - [Googlebot](https://canaicrawl.com/bots/googlebot): Google’s main web crawler. It builds the index behind Google Search, and its fetches also feed AI Overviews unless Google-Extended is disallowed. - [Bingbot](https://canaicrawl.com/bots/bingbot): Microsoft’s crawler for Bing. Its index also powers Copilot answers, DuckDuckGo results and Yahoo Search. - [DuckDuckBot](https://canaicrawl.com/bots/duckduckbot): DuckDuckGo’s crawler. It supplements the Bing index that most DuckDuckGo results come from. - [Applebot](https://canaicrawl.com/bots/applebot): Apple’s crawler for Siri and Spotlight suggestions. When robots.txt doesn’t name it, it follows the Googlebot rules. - [PetalBot](https://canaicrawl.com/bots/petalbot): Huawei’s crawler for Petal Search, the default search on Huawei phones outside China. - [YandexBot](https://canaicrawl.com/bots/yandexbot): The main crawler of Yandex, the largest search engine in Russia. - [Baiduspider](https://canaicrawl.com/bots/baiduspider): The crawler of Baidu, the largest search engine in China. - [OAI-SearchBot](https://canaicrawl.com/bots/oai-searchbot): Builds the index ChatGPT search answers from and links to. It is not used for training; that is GPTBot. - [PerplexityBot](https://canaicrawl.com/bots/perplexitybot): Indexes pages so Perplexity can cite them in answers. Perplexity says it is not used for training. - [Claude-SearchBot](https://canaicrawl.com/bots/claude-searchbot): Indexes pages to improve the quality of Claude’s web search results. Not used for training; that is ClaudeBot. - [DuckAssistBot](https://canaicrawl.com/bots/duckassistbot): Fetches pages for DuckAssist, the AI-generated answers at the top of DuckDuckGo results. - [Amazonbot](https://canaicrawl.com/bots/amazonbot): Amazon’s crawler, used so Alexa and other Amazon services can answer questions from web pages. - [YouBot](https://canaicrawl.com/bots/youbot): The crawler behind You.com’s AI search and answers. - [GPTBot](https://canaicrawl.com/bots/gptbot): Collects pages that may be used to train OpenAI’s models. Blocking it does not affect ChatGPT search, which uses OAI-SearchBot. - [ClaudeBot](https://canaicrawl.com/bots/claudebot): Collects pages that may be used to train Anthropic’s models. Blocking it does not affect Claude’s search or user requests. - [Google-Extended](https://canaicrawl.com/bots/google-extended): Not a crawler. A robots.txt token that tells Google whether pages Googlebot fetched may train Gemini and be used for grounding. Disallowing it does not affect Google Search. - [CCBot](https://canaicrawl.com/bots/ccbot): Builds the free Common Crawl archive, which many AI models have been trained on. - [Applebot-Extended](https://canaicrawl.com/bots/applebot-extended): Not a crawler. A robots.txt token that tells Apple whether pages Applebot fetched may train Apple’s models. Disallowing it does not affect Siri or Spotlight. - [GoogleOther](https://canaicrawl.com/bots/googleother): Google’s generic crawler for research and development crawls by product teams. It is not used for Google Search. - [Meta-ExternalAgent](https://canaicrawl.com/bots/meta-externalagent): Collects pages to train Meta’s AI models and to improve its products. - [Bytespider](https://canaicrawl.com/bots/bytespider): ByteDance’s crawler, believed to collect training data for its models. It has no public documentation and is often reported to ignore robots.txt. - [ChatGPT-User](https://canaicrawl.com/bots/chatgpt-user): Fetches a page when a ChatGPT user asks about it or a GPT action opens it. It does not crawl on its own. - [Claude-User](https://canaicrawl.com/bots/claude-user): Fetches a page when a Claude user asks about it. It does not crawl on its own. - [Perplexity-User](https://canaicrawl.com/bots/perplexity-user): Fetches a page when a Perplexity user asks about it. Perplexity says it generally ignores robots.txt for these requests because a person asked. - [MistralAI-User](https://canaicrawl.com/bots/mistralai-user): Fetches a page when a user of Le Chat asks about it. Mistral says it is not used for training. - [Meta-ExternalFetcher](https://canaicrawl.com/bots/meta-externalfetcher): Fetches a page when a user of Meta AI asks about it. Meta says it may bypass robots.txt because a person asked. - [AhrefsBot](https://canaicrawl.com/bots/ahrefsbot): Builds Ahrefs’ backlink index. Blocking it hides your site from Ahrefs users’ reports, not from search engines. - [SemrushBot](https://canaicrawl.com/bots/semrushbot): Builds Semrush’s backlink and site audit data. Blocking it does not affect search engines. - [MJ12bot](https://canaicrawl.com/bots/mj12bot): Majestic’s distributed crawler for its link index. Blocking it does not affect search engines. - [DotBot](https://canaicrawl.com/bots/dotbot): Moz’s crawler for its link index. Blocking it does not affect search engines.