Blog

GPTBot, ClaudeBot, PerplexityBot: which AI crawlers to block in robots.txt

Every major AI company now crawls the web, and most run several different bots for different jobs. Blocking all of them keeps your content out of AI models, but it can also remove you from AI answers that send you visitors. Blocking none means your pages may be used for training. Here's how to choose.

Three kinds of AI bots

The main user agents

User agentCompanyJob
GPTBotOpenAITraining
OAI-SearchBotOpenAIChatGPT search index
ChatGPT-UserOpenAIFetching for a user
ClaudeBotAnthropicTraining
Claude-SearchBotAnthropicClaude search index
Claude-UserAnthropicFetching for a user
PerplexityBotPerplexitySearch index
Perplexity-UserPerplexityFetching for a user
Google-ExtendedGoogleToken controlling use for Gemini training (crawling is still done by Googlebot)
Applebot-ExtendedAppleToken controlling use for Apple Intelligence training
CCBotCommon CrawlOpen dataset many AI labs train on
meta-externalagentMetaTraining
BytespiderByteDanceTraining

Companies add and rename agents over time, so check each company's documentation before relying on a list, this one included.

Recommended setups

1. Stay visible in AI answers, opt out of training

The most common choice for businesses that want traffic from AI assistants:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: meta-externalagent
Disallow: /

User-agent: Bytespider
Disallow: /

# Search indexers and user agents (OAI-SearchBot, ChatGPT-User,
# Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User)
# stay allowed by the default rule:
User-agent: *
Allow: /

Blocking Google-Extended doesn't affect Google Search rankings. It's a separate token, and regular Search is controlled by Googlebot.

2. Allow everything

If you publish docs, a product or anything you want AI tools to recommend, allowing all agents maximises the chance of being cited. Many SaaS and developer-tool companies do this on purpose.

3. Block everything AI-related

Publishers with paywalled or licensed content often block all of the agents above. Expect to disappear from most AI answers and lose the visits they send.

What robots.txt can't do

Check your current setup

Run your domain through the free AI Crawler Checker. It reads your robots.txt and shows, for each agent above, whether it's allowed, and whether you have an llms.txt file.

Try PageLens on your site

Create a free PageLens account: 7-day trial, no card, no cookie banner needed.

Start free trial

Keep reading