Bot directory · Search engines

Search engines · Baidu

Baiduspider

The crawler of Baidu, the largest search engine in China.

Operator
Baidu
Purpose
Search index. The crawlers behind Google, Bing and the other search engines. Blocking one drops you out of its results.
robots.txt token
Baiduspider
User agent
Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html)
Documentation
www.baidu.com/search/spider.html

Block Baiduspider

Add a group naming it to /robots.txt. Compliant crawlers stop before requesting any page. A firewall rule on the user agent is the only way to enforce it against crawlers that ignore robots.txt.

User-agent: Baiduspider
Disallow: /

Allow Baiduspider when everything else is blocked

A group naming the bot wins over User-agent: *, so an explicit allow lets it in while your catch-all rules stay in place.

User-agent: Baiduspider
Allow: /

User-agent: *
Disallow: /

Does Baiduspider reach your site?

The checker applies your robots.txt the way Baiduspider does, then requests your page with the user agent above, alongside 29 other bots.

Other search engines: Googlebot, Bingbot, DuckDuckBot, Applebot, PetalBot, YandexBot.

Build a complete file with the robots.txt generator.