About canaicrawl
canaicrawl answers one question for any website: which crawlers can actually read it. It is an independent project with no ties to any crawler operator, search engine or SEO vendor.
Why it exists
Whether a site appears in Google, in ChatGPT’s answers or in Perplexity’s citations depends on two things nobody looks at until something breaks: the rules in robots.txt and what the server really does when a crawler shows up. A robots.txt deploy, a new CDN rule or a bot-protection update can lock a crawler out for weeks before anyone notices the traffic is gone. canaicrawl shows both sides in one report, per bot, with the evidence.
How a check works
The baseline request. We load the page you named as ourselves first, with the canaicrawl/1.0 user agent that links to the bot directory and carries a contact address, because that is what a transparent bot should do and what honest bot protection lets through. Only if that is refused do we try again as an ordinary browser, which tells us whether the site is closed to everyone or just to bots. The report says which request got through. Either way this finds where the page really lives after redirects and whether it works at all. If both requests are refused, we mark every live result as unverified and show no score, because a 403 then says nothing about the bot.
robots.txt. We fetch it from the site the page lives on and apply it the way crawlers do, following RFC 9309: the group naming the bot wins, then the User-agent: * group, and the longest matching path rule decides, with Allow winning ties. Applebot follows the Googlebot group when it isn’t named, as Apple documents. A robots.txt that returns a server error counts as “disallow everything”, which is what compliant crawlers do.
A live request per bot. For every bot the rules allow, we request the page again with that bot’s exact user agent and record the status, time and size, and whether a CDN served a challenge page instead. Google-Extended and Applebot-Extended are robots.txt tokens rather than crawlers, so only their rules are checked.
Signals. We read the X-Robots-Tag header and robots meta tags for noindex, nofollow, noai and noimageai, the TDM-Reservation header, Content-Signal lines in robots.txt, llms.txt, and whether a sitemap is declared or present.
The score. You pick a goal: open to everything, search and AI answers without training, or search engines only. Only the bots in that goal count. Groups are weighted (search engines 40, AI search 25, user-triggered fetchers 10, AI training 10, SEO tools 5), share 90 points in proportion to their weight, and earn it by the fraction of their bots that can read the page. llms.txt adds the last 10.
What we send to your site
One check is about 32 requests: the page as canaicrawl (and once more as a browser if that is refused), robots.txt, llms.txt, sitemap.xml when robots.txt lists none, and the page once per bot the rules allow. Our own requests carry canaicrawl and this domain in their user agent; the bot requests carry the user agent listed in the bot directory. We never follow links beyond the page named, submit forms or log in. Results are cached for ten minutes, and new checks are limited per visitor and in how many run at once.
Letting the checker through
Bot-protection services such as Cloudflare verify Googlebot and the other crawlers by the addresses they arrive from, so our requests, which carry those bots’ user agents but come from our own server, are often challenged or refused. When that happens the report says so and shows no score, because a refusal to us proves nothing about the real bot. To get a real result, allow our address or user agent in your firewall or CDN and re-check.
- Address
204.12.218.62- User agent
Mozilla/5.0 (compatible; canaicrawl/1.0; +https://canaicrawl.com/bots; info@canaicrawl.com)for our own requests; the bot directory lists the user agent sent for each bot.- Cloudflare
- A WAF custom rule with the expression
(ip.src eq 204.12.218.62) or (http.user_agent contains "canaicrawl")and the action Skip, placed first. The report shows this with a copy button whenever it is needed.
Limits
Requests come from our server, not from the bot operators’ IP ranges. A firewall that verifies bots by IP can let the real Googlebot in and still refuse us, so a “refused” result is a prompt to check your logs, not a verdict. We check the page you name, not the whole site; rules for other paths may differ. And robots.txt is a request, not a lock: a crawler that ignores it will not show up here as blocked.
Open data
The bots we check are published at /bots.json for scripts and firewalls. Any site can show its score with a badge that links back to its report. The robots.txt generator and llms.txt generator are free to use without an account.
Contact
Corrections to a bot’s user agent or documentation, a crawler we should add, or anything else: info@canaicrawl.com. How we handle data is on the privacy page.