Bot directory · Search engines
Search engines · Baidu
Baiduspider
The crawler of Baidu, the largest search engine in China.
- Operator
- Baidu
- Purpose
- Search index. The crawlers behind Google, Bing and the other search engines. Blocking one drops you out of its results.
- robots.txt token
Baiduspider- User agent
Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html)- Documentation
- www.baidu.com/search/spider.html
Block Baiduspider
Add a group naming it to /robots.txt. Compliant crawlers stop before requesting any page. A firewall rule on the user agent is the only way to enforce it against crawlers that ignore robots.txt.
User-agent: Baiduspider
Disallow: /
Allow Baiduspider when everything else is blocked
A group naming the bot wins over User-agent: *, so an explicit allow lets it in while your catch-all rules stay in place.
User-agent: Baiduspider
Allow: /
User-agent: *
Disallow: /
Does Baiduspider reach your site?
The checker applies your robots.txt the way Baiduspider does, then requests your page with the user agent above, alongside 29 other bots.
Other search engines: Googlebot, Bingbot, DuckDuckBot, Applebot, PetalBot, YandexBot.
Build a complete file with the robots.txt generator.