Free · no signup · robots.txt, live request and noindex

Googlebot checker

Check whether Googlebot can crawl a page on your site: the robots.txt rule Googlebot follows, the response your server gives Googlebot’s user agent, and any noindex that keeps a crawled page out of Google. Then check whether an IP address in your logs is the real Googlebot.

Opens a full report focused on Googlebot, scored for search engines. Bingbot, Applebot, the AI crawlers and SEO bots are checked in the same pass.

What the Googlebot check looks at

01

robots.txt, as Googlebot reads it

We fetch /robots.txt and pick the group Googlebot would: one naming Googlebot, otherwise User-agent: *. The longest matching rule decides, and the report shows its line. A robots.txt that errors counts as blocked, because Google stops crawling then.

02

A live request as Googlebot

We request the page with Googlebot’s user agent and record the status, size and time, whether a CDN answered with a challenge, and which CDN it was. This is how you catch a firewall that refuses crawlers your robots.txt allows.

03

The signals Google obeys

A page can be crawlable and still kept out of Google. We read noindex and nofollow in the robots and googlebot meta tags and the X-Robots-Tag header, and look for a sitemap.

Is it really Googlebot? Verify an IP

Anyone can send Googlebot’s user agent. Paste an address from your access log: we run Google’s reverse-then-forward DNS check and match it against Google’s published IP lists (and Microsoft’s, for Bingbot).

Reading the result

Googlebot resultWhat it meansWhat to do
Allowed · 200Googlebot may crawl the page and your server serves it.If the page is still missing from Google, the cause is indexing, not access: check for noindex under Signals, then the page’s status in Search Console.
Blocked · robots.txtA Disallow rule matches Googlebot. The report shows the line.Remove the rule, or add a narrower Allow. A User-agent: * block catches Googlebot unless it has its own group.
403, 429 or a challengeYour server, firewall or CDN refused Googlebot’s user agent.See the Cloudflare guide. If the report says our checker itself was refused, the real Googlebot may still get through; check your CDN’s security events.
5xx or a timeoutThe server failed. Google slows down on errors, and an error on robots.txt stops crawling.Fix the error and re-check. Keep /robots.txt fast and outside any bot protection.
noindex under SignalsGooglebot can fetch the page but is told not to index it.Remove the meta tag or X-Robots-Tag header if the page should be in Google. Don’t also disallow it, or Google can’t see the noindex.

Other ways to check Googlebot access

Why our check can differ from Google’s

We send Googlebot’s user agent from our own server. Bot protection that verifies Googlebot by its IP addresses may refuse us while letting Google in; the report then says the checker was refused and marks bots unverified instead of blocked. The reverse can happen too: a site can serve the real Googlebot something different from what it serves us. Googlebot also crawls mostly from the United States, so a country block can stop Google while our check passes. When the two disagree, Search Console is the authority on what Google saw.

For the exact user agent strings Googlebot sends, and the rest of Google’s crawlers, see Googlebot user agent. For Bing, see Bingbot.

Questions

How do I check if Googlebot can crawl my site?

Enter a page above. The checker reads your robots.txt and applies the group Googlebot would use, requests the page with Googlebot’s user agent, and reads the noindex signals Google obeys. It then opens the report on Googlebot, with the matching robots.txt line and the HTTP response. For the definitive answer from Google’s own addresses, use URL Inspection’s live test in Search Console.

Why would Googlebot be blocked on my site?

The usual causes are a robots.txt Disallow that matches Googlebot (often a catch-all User-agent: * group meant for other bots), a robots.txt that returns a server error, which makes Google stop crawling, a CDN or firewall challenge that crawlers can’t solve, a firewall rule on the word “bot”, and a country block that includes the United States, where most Googlebot requests come from.

How do I know if a visitor claiming to be Googlebot is real?

Check its IP address, not its user agent. A real Googlebot address has reverse DNS ending in googlebot.com or google.com, the hostname resolves back to the same address, and the address is in Google’s common-crawlers.json list. Paste the address into the verifier on this page to run those checks.

Can I see my page the way Googlebot sees it?

URL Inspection in Search Console fetches and renders the page from Google’s own addresses and shows the result, for sites you have verified. This checker shows what your server answers to Googlebot’s user agent from outside Google, which is what catches firewall and CDN rules.

Is the result here the same as what Google gets?

Usually, but not always. Our requests carry Googlebot’s user agent but come from our own server, not Google’s addresses. Bot protection that verifies Googlebot by IP may refuse us and let Google in; the report then marks the result unverified rather than blocked. Your CDN’s security events or Search Console’s Crawl stats show what the real Googlebot got.

Is this Google bot checker free?

Yes. There is no signup, and a check of any public page takes a few seconds. The same report covers Bingbot, the AI crawlers and SEO bots.