What are AI crawlers?
AI crawlers are the automated user agents that AI companies use to fetch web pages. They differ from Googlebot in what they do with the pages afterwards. OpenAI’s documentation draws the same three-way split that other vendors use: some bots collect content that may be used to train models, some build the index that powers AI search, and some visit a page only because a user asked a question.
- Training: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, Meta-ExternalAgent. They collect content that may be used to train models, or in the case of Google-Extended and Applebot-Extended, they are control tokens that govern that use.
- AI search: OAI-SearchBot, Claude-SearchBot, PerplexityBot, MistralAI-Index. They build the indexes that AI search products cite. Blocking them can remove you from that product’s answers.
- User fetch: ChatGPT-User, Claude-User, Perplexity-User. They fetch a page on request, for example when someone asks a question that needs it.
Which AI crawlers should you know?
The table lists every crawler Serpel checks, with what each vendor’s own documentation says. The purpose column is Serpel’s classification. CCBot, for example, counts as training because Common Crawl’s open data can be used by any organisation.
| User agent | Vendor | Purpose | What the vendor says |
|---|---|---|---|
| OAI-SearchBot | OpenAI | AI search | Surfaces websites in ChatGPT’s search features. Sites that opt out are not shown in ChatGPT search answers, and changes can take about 24 hours to apply. OpenAI docs |
| GPTBot | OpenAI | Training | Crawls content that may be used to train OpenAI’s generative AI foundation models. Disallowing it signals that your content should not be used for training. OpenAI docs |
| ChatGPT-User | OpenAI | User fetch | Visits a page when a user asks ChatGPT or a custom GPT something that needs it. OpenAI says robots.txt rules may not apply to these user-initiated requests. OpenAI docs |
| Claude-SearchBot | Anthropic | AI search | Analyses content to improve the relevance of Claude’s search results. Anthropic says disabling it may reduce your visibility in search results. Anthropic docs |
| ClaudeBot | Anthropic | Training | Collects web content that could contribute to model training. A block tells Anthropic to exclude your future materials from training datasets. Anthropic docs |
| Claude-User | Anthropic | User fetch | Accesses pages when a Claude user asks a question that needs them. Anthropic says disabling it may reduce your visibility in user-directed search. Anthropic docs |
| PerplexityBot | Perplexity | AI search | Surfaces and links websites in Perplexity search results. Perplexity says it is not used to crawl content for AI foundation models. Perplexity docs |
| Perplexity-User | Perplexity | User fetch | Visits a page to answer a user’s question and links to it. Perplexity says this fetcher generally ignores robots.txt because a user requested the fetch. Perplexity docs |
| Google-Extended | Training | A robots.txt token rather than a separate crawler. It controls whether content Google crawls can be used for Gemini training and grounding, and Google says it does not affect inclusion or ranking in Search. Google docs | |
| Applebot-Extended | Apple | Training | Lets you opt out of Apple using your content to train its foundation models. Apple says it does not crawl pages itself and that blocked pages can still appear in search results. Apple docs |
| CCBot | Common Crawl | Training | The crawler of Common Crawl, a non-profit that publishes open web crawl data for any organisation to use. A Disallow rule keeps your pages out of its future crawls. Common Crawl docs |
| Meta-ExternalAgent | Meta | Training | Crawls to train AI models or to improve products by indexing content directly, according to Meta. Crawlers may cache robots.txt for up to 24 hours. Meta docs |
| MistralAI-Index | Mistral | AI search | Crawls for indexing only, to power Mistral’s search. Mistral says content it crawls is not used for generative AI training. Mistral docs |
Should you block AI crawlers or allow them?
It depends on what you want from AI products. The groups are independent, so you can mix them. OpenAI states that a site can allow OAI-SearchBot to appear in search while disallowing GPTBot so its content is not used for training.
| Goal | Allow | Block |
|---|---|---|
| Appear in AI answers, opt out of training | OAI-SearchBot, Claude-SearchBot, PerplexityBot, MistralAI-Index, ChatGPT-User, Claude-User, Perplexity-User | GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, Meta-ExternalAgent |
| Appear everywhere, including training datasets | All of them | None |
| Opt out of every AI crawler | None | OAI-SearchBot, GPTBot, ChatGPT-User, Claude-SearchBot, ClaudeBot, Claude-User, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, CCBot, Meta-ExternalAgent, MistralAI-Index |
Two points often cause confusion. First, Google’s AI Overviews and AI Mode draw on the normal Search index, so they are controlled with Googlebot rules and snippet controls, not with Google-Extended. Google says robots.txt rules for Googlebot govern crawling for Search including AI features, and that nosnippet, data-nosnippet, max-snippet and noindex limit what is shown. Blocking Googlebot would remove you from Google Search, so never use it to opt out of AI. Second, a block affects future crawling. Anthropic describes a ClaudeBot block as excluding a site’s future materials from training datasets, so do not expect robots.txt to undo collection that already happened.
Ready-to-copy robots.txt examples
The files below are generated from the same crawler list the table uses. For a GPTBot robots.txt rule or a ClaudeBot robots.txt rule, add the crawler’s name as a User-agent line. Replace the example paths with your own and keep your existing rules for other crawlers. Because a crawler follows the group that names it and ignores the general group, the allowed bots repeat the Disallow line for the private path.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: MistralAI-Index
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
Allow: /
Disallow: /admin/User-agent: OAI-SearchBot
User-agent: GPTBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: ClaudeBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Meta-ExternalAgent
User-agent: MistralAI-Index
Disallow: /User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Meta-ExternalAgent
Disallow: /members/
Disallow: /downloads/Several user-agent lines above one set of rules form a single group, which RFC 9309 allows. Product tokens are matched case-insensitively, so GPTBot and gptbot are the same rule. Search systems need time to notice a change: OpenAI says about 24 hours, and Meta says its crawlers may cache robots.txt for up to 24 hours.
Does robots.txt actually stop AI crawlers?
Only the ones that choose to obey it. RFC 9309 is explicit that the rules are not a form of access authorisation. Vendors document what their bots do:
- Training and search crawlers. Anthropic says its bots honour robots.txt, OpenAI manages OAI-SearchBot and GPTBot through it, and Meta presents it as the way to block Meta-ExternalAgent.
- User-triggered fetchers. OpenAI says robots.txt rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores them, and Meta says Meta-ExternalFetcher may bypass them.
- Verification. OpenAI, Perplexity, Anthropic, Common Crawl and Mistral publish IP ranges for their crawlers, so you can tell a real request from one that merely copies a name. Anthropic warns that blocking by IP address may not reliably guarantee an opt-out, so it recommends robots.txt.
If you need a hard block, use a firewall rule or authentication. The reverse problem is just as common: bot protection can block a crawler that robots.txt allows. OpenAI recommends allowing its published IP ranges for OAI-SearchBot, and Vercel’s analysis of AI crawlers mentions a firewall rule that blocks AI bots, which would also stop the ones you want.
How do you check which AI crawlers your site allows?
Test your robots.txt
Paste your domain into the free robots.txt checker to see which AI crawlers your file allows and which it blocks.
Count the requests in your logs
The command below counts requests per AI crawler in an access log. Google-Extended and Applebot-Extended never appear, because Google says Google-Extended has no separate user agent and Apple says Applebot-Extended does not crawl pages.
Count AI crawler requests grep -o -i -E "OAI-SearchBot|GPTBot|ChatGPT-User|Claude-SearchBot|ClaudeBot|Claude-User|PerplexityBot|Perplexity-User|CCBot|Meta-ExternalAgent|MistralAI-Index" access.log | sort | uniq -c | sort -rnAudit it with every crawl
Serpel’s site audit raises a warning when robots.txt blocks an AI search crawler, and
serpel ai statuslists which AI crawlers your robots.txt allows from the last crawl. AI visibility in Serpel also shows whether answers cite you.
Allowing the search crawlers is only the first step towards being cited. Our guide to SEO for ChatGPT covers the rest, and the llms.txt examples explain why llms.txt is a different file with a different job.
Frequently asked questions
Should I block GPTBot?
Block GPTBot if you do not want OpenAI to use your content for training. It has no effect on ChatGPT search, which uses OAI-SearchBot, and OpenAI says each setting is independent. If you want to be cited in ChatGPT answers, keep OAI-SearchBot allowed.
Does blocking AI crawlers hurt my Google rankings?
Blocking the AI-specific tokens does not. Google says Google-Extended does not affect inclusion or ranking in Google Search, and GPTBot and ClaudeBot are separate from search crawlers. Blocking Googlebot would remove you from Google Search, so never use it to opt out of AI.
How do I block all AI crawlers?
Add one group to robots.txt with a User-agent line for each crawler and Disallow: / underneath. The second example above does this for all 13 crawlers Serpel checks. Robots.txt is a request, not enforcement, so user-triggered fetchers may still visit and a hard block needs a firewall rule or authentication.
Can AI crawlers ignore robots.txt?
Yes. Under RFC 9309 robots.txt is voluntary and not a form of access authorisation. Vendors say their training and search crawlers honour it, but OpenAI, Perplexity and Meta document that user-initiated fetchers such as ChatGPT-User, Perplexity-User and Meta-ExternalFetcher may not.
What is Google-Extended?
Google-Extended is a robots.txt token that controls whether content Google crawls can be used for Gemini training and grounding. It has no separate user agent, and Google says it does not affect inclusion or ranking in Google Search.
Sources
- OpenAI: Overview of OpenAI crawlers, accessed 10 Oct 2026
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?, accessed 10 Oct 2026
- Perplexity: Perplexity crawlers, accessed 10 Oct 2026
- Google Search Central: Google’s common crawlers (Google-Extended), accessed 10 Oct 2026
- Google Search Central: AI features and your website, accessed 10 Oct 2026
- Apple: About Applebot, accessed 10 Oct 2026
- Common Crawl: CCBot, accessed 10 Oct 2026
- Meta: Web crawlers (Meta-ExternalAgent, Meta-ExternalFetcher), accessed 10 Oct 2026
- Mistral AI: Robots (MistralAI-Index), accessed 10 Oct 2026
- RFC 9309: Robots Exclusion Protocol, accessed 10 Oct 2026
- Vercel: The rise of the AI crawler, accessed 10 Oct 2026
Related reading
- Robots.txt checkerFree robots.txt checker: see which of 13 AI crawlers can reach your homepage, list your sitemaps and check llms.txt.
- SEO for ChatGPT: how to rank in ChatGPT search and get citedSEO for ChatGPT means letting OAI-SearchBot in and being quotable. How ChatGPT search picks sources, what to do and how to check citations.
- llms.txt example: 3 complete files for docs, shops and blogsThree complete llms.txt examples for a SaaS docs site, an online shop and a blog, plus the format rules, llms-full.txt and a validation script.
- Generative engine optimization (GEO): what it is and what worksGenerative engine optimization (GEO) is how you get cited in AI answers. See what the research found, what Google says and a checklist you can use.
