# Robots.txt checker for AI crawlers

URL: https://serpel.app/tools/robots-txt-checker

Updated: 2026-10-10

A robots.txt checker fetches a domain’s robots.txt and shows what the file allows for each crawler. This free tool tests it against all 13 AI crawlers Serpel knows, shows which of them can reach your homepage, lists your sitemaps and checks for llms.txt and llms-full.txt.

## What does a robots.txt checker test?

A robots.txt checker fetches the `robots.txt` file of a domain and applies its rules the way a crawler would. This one tests all 13 AI crawlers Serpel knows and shows which of them can reach your homepage.

It also lists the sitemaps declared in the file, which the [sitemap checker](https://serpel.app/tools/sitemap-checker) can validate, and checks whether `llms.txt` and `llms-full.txt` exist at the root of your domain. If the file is missing, the [llms.txt generator](https://serpel.app/tools/llms-txt-generator) builds one. Use the tool as a robots.txt tester after every change, or to check robots.txt before you launch a site.

A crawler counts as allowed when the group that applies to it does not disallow the homepage. Check that every result matches your intent.

## How does robots.txt work?

The Robots Exclusion Protocol is standardised in [RFC 9309](https://www.rfc-editor.org/rfc/rfc9309.html). Google describes its own reading in the [robots.txt specification](https://developers.google.com/crawling/docs/robots-txt/robots-txt-spec). These are the rules that matter most:

- **Location.** The file must be served at `/robots.txt` in the top-level path of a host. Its rules apply only to that protocol, host and port, so every subdomain needs its own file.
- **Groups.** A group starts with one or more `User-agent` lines, followed by `Allow` and `Disallow` rules. A crawler matches its product token case-insensitively and obeys only the matching group. If none matches, it falls back to the `*` group. Specific groups and the `*` group are never merged.
- **Longest match wins.** The most specific rule, meaning the one with the longest matching path, applies. If an `Allow` and a `Disallow` rule are equally specific, RFC 9309 says the `Allow` rule should win. Google applies the least restrictive rule in a conflict.
- **Status codes.** A `2xx` response is parsed. A `4xx` response means there is no file, so crawlers may fetch anything. A `5xx` response or a timeout means the file is unreachable, and RFC 9309 tells crawlers to assume that everything is disallowed. Google pauses crawling, keeps retrying and falls back to its last cached copy for up to 30 days.
- **Size and caching.** Crawlers must parse at least 500 KiB. Google generally caches the file for up to 24 hours, so a change can take a day to reach every bot.
- **Not access control.** RFC 9309 states that these rules are not a form of access authorisation. Private content needs authentication, not a `Disallow` line.

## Which AI crawlers should you know about?

AI companies run crawlers for three different purposes, and you can treat each purpose separately:

- **Training** crawlers collect content that may be used to train foundation models.
- **AI search** crawlers index pages so that an assistant can cite and link them in its answers.
- **User fetch** agents visit a single page because a person asked an assistant about it.

**The AI crawlers Serpel checks**
| User agent | Vendor | Purpose |
| --- | --- | --- |
| `OAI-SearchBot` | OpenAI | AI search |
| `GPTBot` | OpenAI | Training |
| `ChatGPT-User` | OpenAI | User fetch |
| `Claude-SearchBot` | Anthropic | AI search |
| `ClaudeBot` | Anthropic | Training |
| `Claude-User` | Anthropic | User fetch |
| `PerplexityBot` | Perplexity | AI search |
| `Perplexity-User` | Perplexity | User fetch |
| `Google-Extended` | Google | Training |
| `Applebot-Extended` | Apple | Training |
| `CCBot` | Common Crawl | Training |
| `Meta-ExternalAgent` | Meta | Training |
| `MistralAI-Index` | Mistral | AI search |

Three details are easy to miss:

- **Control tokens.** `Google-Extended` and `Applebot-Extended` do not crawl on their own. Google says [Google-Extended](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers) has no separate user agent and does not affect inclusion in Google Search. Apple says [Applebot-Extended](https://support.apple.com/en-us/119829) only decides whether content already crawled by Applebot may be used for training.
- **User fetches.** [OpenAI](https://developers.openai.com/api/docs/bots) says `ChatGPT-User` may not follow robots.txt, and [Perplexity](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) says `Perplexity-User` generally ignores it. [Anthropic](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) says all of its bots honour robots.txt.
- **Opt-outs look forward.** Blocking a training crawler is a request about future collection. Anthropic, for example, describes it as a signal that future content should be excluded.

## How do you allow AI search but block training?

Give each training crawler a `Disallow: /` rule and leave the AI search crawlers alone. They then follow your `*` group, which allows everything except the paths you list. This example blocks the 6 training crawlers Serpel knows:

```text
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Meta-ExternalAgent
Disallow: /

User-agent: *
Disallow: /account/

Sitemap: https://example.com/sitemap.xml
```

`OAI-SearchBot`, `Claude-SearchBot`, `PerplexityBot` and `MistralAI-Index` need no group of their own. If you add one, repeat the rules from your `*` group inside it, because a crawler obeys only one group. Run the checker again to confirm the result. Which setup is right depends on your goals, and our [guide to AI crawlers](https://serpel.app/blog/ai-crawlers) explains the trade-offs.

## What are the most common robots.txt mistakes?

- **A leftover blanket block.** A staging file with `User-agent: *` and `Disallow: /` that reaches production blocks every crawler, search engines included.
- **A bot group that replaces the wildcard group.** Adding a group for one bot with a single rule means that bot no longer sees the rules from your `*` group.
- **Hiding pages from search with Disallow.** robots.txt controls crawling, not indexing. As Google explains, a blocked URL can still be indexed and shown without a snippet.
- **Errors on the file itself.** A `5xx` response or a timeout on `/robots.txt` can stop crawling altogether. A bot-protection rule that answers crawlers with a challenge page blocks them whatever the file says.
- **Typos in tokens.** `GPT-Bot` is not `GPTBot`, and a misspelt token matches nothing.

## What does this robots.txt checker not do?

- It tests the homepage path `/` only. A rule such as `Disallow: /blog/` will not show up as a block.
- It fetches only `robots.txt`, `llms.txt` and `llms-full.txt`. It does not crawl pages, render JavaScript or read `noindex` tags and `X-Robots-Tag` headers.
- It does not request your pages as those crawlers. A firewall or CDN rule can still block them, and only your server logs show that.
- It does not lint the file line by line, so it is not a strict robots.txt validator.

For the whole site, Serpel’s [AI visibility features](https://serpel.app/features/ai-visibility) check crawler access and llms.txt on every crawl and show whether ChatGPT and Google AI Overviews cite you.

## Frequently asked questions

### How do I check whether GPTBot can crawl my site?

Enter your domain in the checker and look for GPTBot in the results. It counts as allowed when the robots.txt group that applies to GPTBot does not disallow your homepage. The checker only reads robots.txt, so a firewall or bot-protection rule can still block the crawler.

### Does blocking GPTBot remove my site from ChatGPT?

No. GPTBot is OpenAI’s training crawler, while ChatGPT search uses OAI-SearchBot. [OpenAI says](https://developers.openai.com/api/docs/bots) that sites opted out of OAI-SearchBot will not appear in ChatGPT search answers. Block GPTBot to opt out of training and leave OAI-SearchBot allowed to stay visible.

### What happens if my robots.txt returns a 404 or a server error?

A 4xx response such as 404 means there is no file, so crawlers may fetch everything. A 5xx response or a timeout counts as unreachable, and [RFC 9309](https://www.rfc-editor.org/rfc/rfc9309.html) tells crawlers to assume that everything is disallowed. Google pauses crawling and falls back to its last cached copy for up to 30 days.

### Do AI crawlers have to obey robots.txt?

robots.txt is a convention that crawlers choose to follow, and OpenAI, Anthropic and Perplexity document that their crawlers do. The exceptions are fetches a user triggers: OpenAI says ChatGPT-User may not follow robots.txt, and Perplexity says Perplexity-User generally ignores it. For anything private, use authentication instead.

### Does the checker test llms.txt too?

Yes. It reports whether llms.txt and llms-full.txt exist at the root of the domain. If you need the file, the [llms.txt generator](https://serpel.app/tools/llms-txt-generator) creates it.