Free tool

Robots.txt checker for AI crawlers

A robots.txt checker fetches a domain’s robots.txt and shows what the file allows for each crawler. This free tool tests it against all 13 AI crawlers Serpel knows, shows which of them can reach your homepage, lists your sitemaps and checks for llms.txt and llms-full.txt.

Free, no sign-up. We read the public robots.txt and llms.txt of the site and don’t store the result or your input.

What does a robots.txt checker test?

A robots.txt checker fetches the robots.txt file of a domain and applies its rules the way a crawler would. This one tests all 13 AI crawlers Serpel knows and shows which of them can reach your homepage.

It also lists the sitemaps declared in the file, which the sitemap checker can validate, and checks whether llms.txt and llms-full.txt exist at the root of your domain. If the file is missing, the llms.txt generator builds one. Use the tool as a robots.txt tester after every change, or to check robots.txt before you launch a site.

A crawler counts as allowed when the group that applies to it does not disallow the homepage. Check that every result matches your intent.

How does robots.txt work?

The Robots Exclusion Protocol is standardised in RFC 9309. Google describes its own reading in the robots.txt specification. These are the rules that matter most:

  • Location. The file must be served at /robots.txt in the top-level path of a host. Its rules apply only to that protocol, host and port, so every subdomain needs its own file.
  • Groups. A group starts with one or more User-agent lines, followed by Allow and Disallow rules. A crawler matches its product token case-insensitively and obeys only the matching group. If none matches, it falls back to the * group. Specific groups and the * group are never merged.
  • Longest match wins. The most specific rule, meaning the one with the longest matching path, applies. If an Allow and a Disallow rule are equally specific, RFC 9309 says the Allow rule should win. Google applies the least restrictive rule in a conflict.
  • Status codes. A 2xx response is parsed. A 4xx response means there is no file, so crawlers may fetch anything. A 5xx response or a timeout means the file is unreachable, and RFC 9309 tells crawlers to assume that everything is disallowed. Google pauses crawling, keeps retrying and falls back to its last cached copy for up to 30 days.
  • Size and caching. Crawlers must parse at least 500 KiB. Google generally caches the file for up to 24 hours, so a change can take a day to reach every bot.
  • Not access control. RFC 9309 states that these rules are not a form of access authorisation. Private content needs authentication, not a Disallow line.

Which AI crawlers should you know about?

AI companies run crawlers for three different purposes, and you can treat each purpose separately:

  • Training crawlers collect content that may be used to train foundation models.
  • AI search crawlers index pages so that an assistant can cite and link them in its answers.
  • User fetch agents visit a single page because a person asked an assistant about it.
The AI crawlers Serpel checks
User agentVendorPurpose
OAI-SearchBotOpenAIAI search
GPTBotOpenAITraining
ChatGPT-UserOpenAIUser fetch
Claude-SearchBotAnthropicAI search
ClaudeBotAnthropicTraining
Claude-UserAnthropicUser fetch
PerplexityBotPerplexityAI search
Perplexity-UserPerplexityUser fetch
Google-ExtendedGoogleTraining
Applebot-ExtendedAppleTraining
CCBotCommon CrawlTraining
Meta-ExternalAgentMetaTraining
MistralAI-IndexMistralAI search

Three details are easy to miss:

  • Control tokens. Google-Extended and Applebot-Extended do not crawl on their own. Google says Google-Extended has no separate user agent and does not affect inclusion in Google Search. Apple says Applebot-Extended only decides whether content already crawled by Applebot may be used for training.
  • User fetches. OpenAI says ChatGPT-User may not follow robots.txt, and Perplexity says Perplexity-User generally ignores it. Anthropic says all of its bots honour robots.txt.
  • Opt-outs look forward. Blocking a training crawler is a request about future collection. Anthropic, for example, describes it as a signal that future content should be excluded.

How do you allow AI search but block training?

Give each training crawler a Disallow: / rule and leave the AI search crawlers alone. They then follow your * group, which allows everything except the paths you list. This example blocks the 6 training crawlers Serpel knows:

robots.txt
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Meta-ExternalAgent
Disallow: /

User-agent: *
Disallow: /account/

Sitemap: https://example.com/sitemap.xml

OAI-SearchBot, Claude-SearchBot, PerplexityBot and MistralAI-Index need no group of their own. If you add one, repeat the rules from your * group inside it, because a crawler obeys only one group. Run the checker again to confirm the result. Which setup is right depends on your goals, and our guide to AI crawlers explains the trade-offs.

What are the most common robots.txt mistakes?

  • A leftover blanket block. A staging file with User-agent: * and Disallow: / that reaches production blocks every crawler, search engines included.
  • A bot group that replaces the wildcard group. Adding a group for one bot with a single rule means that bot no longer sees the rules from your * group.
  • Hiding pages from search with Disallow. robots.txt controls crawling, not indexing. As Google explains, a blocked URL can still be indexed and shown without a snippet.
  • Errors on the file itself. A 5xx response or a timeout on /robots.txt can stop crawling altogether. A bot-protection rule that answers crawlers with a challenge page blocks them whatever the file says.
  • Typos in tokens. GPT-Bot is not GPTBot, and a misspelt token matches nothing.

What does this robots.txt checker not do?

  • It tests the homepage path / only. A rule such as Disallow: /blog/ will not show up as a block.
  • It fetches only robots.txt, llms.txt and llms-full.txt. It does not crawl pages, render JavaScript or read noindex tags and X-Robots-Tag headers.
  • It does not request your pages as those crawlers. A firewall or CDN rule can still block them, and only your server logs show that.
  • It does not lint the file line by line, so it is not a strict robots.txt validator.

For the whole site, Serpel’s AI visibility features check crawler access and llms.txt on every crawl and show whether ChatGPT and Google AI Overviews cite you.

Frequently asked questions

How do I check whether GPTBot can crawl my site?

Enter your domain in the checker and look for GPTBot in the results. It counts as allowed when the robots.txt group that applies to GPTBot does not disallow your homepage. The checker only reads robots.txt, so a firewall or bot-protection rule can still block the crawler.

Does blocking GPTBot remove my site from ChatGPT?

No. GPTBot is OpenAI’s training crawler, while ChatGPT search uses OAI-SearchBot. OpenAI says that sites opted out of OAI-SearchBot will not appear in ChatGPT search answers. Block GPTBot to opt out of training and leave OAI-SearchBot allowed to stay visible.

What happens if my robots.txt returns a 404 or a server error?

A 4xx response such as 404 means there is no file, so crawlers may fetch everything. A 5xx response or a timeout counts as unreachable, and RFC 9309 tells crawlers to assume that everything is disallowed. Google pauses crawling and falls back to its last cached copy for up to 30 days.

Do AI crawlers have to obey robots.txt?

robots.txt is a convention that crawlers choose to follow, and OpenAI, Anthropic and Perplexity document that their crawlers do. The exceptions are fetches a user triggers: OpenAI says ChatGPT-User may not follow robots.txt, and Perplexity says Perplexity-User generally ignores it. For anything private, use authentication instead.

Does the checker test llms.txt too?

Yes. It reports whether llms.txt and llms-full.txt exist at the root of the domain. If you need the file, the llms.txt generator creates it.

Updated 10 Oct 2026

Keep reading