GPTBot, ClaudeBot, PerplexityBot: which AI crawlers to block in robots.txt
Every major AI company now crawls the web, and most run several different bots for different jobs. Blocking all of them keeps your content out of AI models, but it can also remove you from AI answers that send you visitors. Blocking none means your pages may be used for training. Here's how to choose.
Three kinds of AI bots
- Training crawlers collect pages to train future models. Blocking them doesn't change what assistants show today.
- Search indexers build the index that an assistant's search feature uses to find and cite sources. Blocking them makes it less likely you're cited, and so less likely you get clicks.
- User agents fetch a page when a person asks the assistant to read it ("summarise this URL") or while it browses for an answer. They act on a user's request and work more like a browser than a crawler.
The main user agents
| User agent | Company | Job |
|---|---|---|
GPTBot | OpenAI | Training |
OAI-SearchBot | OpenAI | ChatGPT search index |
ChatGPT-User | OpenAI | Fetching for a user |
ClaudeBot | Anthropic | Training |
Claude-SearchBot | Anthropic | Claude search index |
Claude-User | Anthropic | Fetching for a user |
PerplexityBot | Perplexity | Search index |
Perplexity-User | Perplexity | Fetching for a user |
Google-Extended | Token controlling use for Gemini training (crawling is still done by Googlebot) | |
Applebot-Extended | Apple | Token controlling use for Apple Intelligence training |
CCBot | Common Crawl | Open dataset many AI labs train on |
meta-externalagent | Meta | Training |
Bytespider | ByteDance | Training |
Companies add and rename agents over time, so check each company's documentation before relying on a list, this one included.
Recommended setups
1. Stay visible in AI answers, opt out of training
The most common choice for businesses that want traffic from AI assistants:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: meta-externalagent
Disallow: /
User-agent: Bytespider
Disallow: /
# Search indexers and user agents (OAI-SearchBot, ChatGPT-User,
# Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User)
# stay allowed by the default rule:
User-agent: *
Allow: /
Blocking Google-Extended doesn't affect Google Search rankings. It's a separate token, and regular Search is controlled by Googlebot.
2. Allow everything
If you publish docs, a product or anything you want AI tools to recommend, allowing all agents maximises the chance of being cited. Many SaaS and developer-tool companies do this on purpose.
3. Block everything AI-related
Publishers with paywalled or licensed content often block all of the agents above. Expect to disappear from most AI answers and lose the visits they send.
What robots.txt can't do
- It's a request, not a lock. Reputable crawlers follow it; others may not. To enforce a block, filter by user agent or IP range on your server or CDN. Cloudflare, for example, has a one-click AI-bot block.
- It doesn't remove content already collected. It only affects future crawls.
- It tells you who may visit, not who does. To see which AI bots actually read your pages and how many people AI assistants send you, you need analytics that classifies bots, such as PageLens AI traffic analytics.
Check your current setup
Run your domain through the free AI Crawler Checker. It reads your robots.txt and shows, for each agent above, whether it's allowed, and whether you have an llms.txt file.