What is LLM SEO?
LLM SEO is search engine optimisation for the assistants that sit on top of large language models (LLMs). When someone asks ChatGPT, Claude, Gemini or Perplexity a question, the answer comes from the model’s training, from pages the assistant retrieves while it answers, or from both. LLM SEO is the work of making sure your site is in the second group, readable by the bots that retrieve it, and worth quoting.
You will also hear it called “SEO for AI”, AI SEO, GEO and AEO. The labels overlap, and Google’s optimisation guide says optimising for generative AI search is still SEO. Our GEO vs SEO comparison explains how the terms relate, and the AI SEO guide covers using AI as a tool for SEO work. This article is about being found and cited by LLM assistants.
How do LLM assistants find and cite sources?
An LLM on its own is not a search engine. It generates text from patterns it learned in training. To cite a page, the product around the model has to retrieve that page at answer time. The vendors describe this in their documentation.
- ChatGPT: OpenAI says ChatGPT may search the web automatically when a question benefits from current information, typically rewrites the question into targeted queries and may attach citations. It also warns that results and citations can be incomplete, outdated or incorrect.
- Claude: Anthropic’s web search tool lets Claude decide when to search, run the searches and answer with cited sources. Claude answers directly for stable knowledge and searches for current or changing information. The page does not say which index it searches.
- Gemini: Google’s grounding with Google Search lets the model decide whether a search helps, generate queries, run them and answer with citations to the sources.
- Perplexity: its crawler documentation says PerplexityBot is designed to surface and link websites in Perplexity search results, and that Perplexity-User visits a page to answer a user’s question and includes a link.
Training data versus live retrieval
| Path | How it works | What you can do | When it takes effect |
|---|---|---|---|
| Training data | Crawlers collect web content that may be used to train future models. What a released model learned is fixed | Allow or disallow training bots such as GPTBot and ClaudeBot in robots.txt, which is the control the vendors document | Only for future model versions. A block applies to future crawls, not to what a released model already learned |
| Live retrieval | The assistant searches or fetches pages while it answers and cites them | Allow the search bots, stay indexed, serve content as text and write quotable passages | As soon as crawlers revisit. OpenAI says search systems can take about 24 hours to adjust to a robots.txt change |
| User-requested fetch | A user asks about a specific page and the assistant fetches it | Make sure the page loads quickly and returns status 200 to the user-triggered bots | Immediately, on each request |
The practical consequence is that a citation needs retrieval. If an assistant answers from memory it will not link to you, and what it says about your product depends on what was written about it before its training cut-off. That is why most LLM SEO effort goes into retrieval.
Which crawlers should you allow for LLM SEO?
Each vendor splits its bots by purpose, so you can allow search and block training. This table shows the documented behaviour. Our guide to AI crawlers lists the rest.
| Vendor | Bot | Purpose | robots.txt |
|---|---|---|---|
| OpenAI | OAI-SearchBot | Surfaces websites in ChatGPT’s search features | Applies. Sites that opt out are not shown in ChatGPT search answers |
| OpenAI | GPTBot | Crawls content that may be used to train foundation models | Applies. Disallow it to opt out of training |
| OpenAI | ChatGPT-User | Visits pages for certain user actions | OpenAI says the rules may not apply to user-initiated requests |
| Anthropic | Claude-SearchBot | Improves the quality of Claude’s search results | Honours robots.txt. Disabling it may reduce visibility in search results |
| Anthropic | ClaudeBot | Collects content that could contribute to model training | Honours robots.txt. A block excludes future material from training |
| Anthropic | Claude-User | Retrieves pages when a user asks a question | Honours robots.txt. Disabling it may reduce visibility for user-directed search |
| Perplexity | PerplexityBot | Surfaces and links websites in Perplexity results. Not used to train foundation models | Applies. Perplexity recommends allowing it |
| Perplexity | Perplexity-User | Visits a page to answer a user’s question | Generally ignores robots.txt, because a user requested the fetch |
| Google-Extended | A robots.txt token that controls use of crawled content for Gemini training and grounding | A token only, with no separate user agent. Does not affect inclusion in Search or ranking |
Sources: the documentation of OpenAI, Anthropic, Perplexity and Google. OpenAI and Perplexity each say their settings work independently, and Anthropic documents its three bots separately, so the file below allows search and user requests and opts out of training.
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
Allow: /
User-agent: GPTBot
User-agent: ClaudeBot
Disallow: /Google-Extended is left out on purpose. Google says the token covers both Gemini training and grounding, which means serving the Search index to the model at prompt time, so decide on it separately. Whether to opt out of training at all is a business decision, not an SEO one, and it does not decide whether you can be cited in live answers. Allowing the crawler is not enough on its own. A firewall or bot protection can block a bot that robots.txt allows, so check your CDN rules against the vendors’ published IP ranges, such as OpenAI’s OAI-SearchBot list and Anthropic’s list. To see how your file reads, use the free robots.txt checker. It tests your robots.txt against 13 AI crawlers, as every Serpel crawl does.
Do LLM crawlers run JavaScript?
Mostly not, according to the best public evidence. In a December 2024 analysis of its network, Vercel found that none of the major AI crawlers it measured rendered JavaScript, with the exceptions of Gemini, which uses Googlebot’s infrastructure, and Applebot. ChatGPT and Claude fetched JavaScript files (11.50% and 23.84% of their requests) but did not execute them. The data is almost two years old, comes from Vercel’s own network and predates later crawler changes, so read it as a strong hint, not a current guarantee.
The safe rule is to put everything an assistant should quote in the initial HTML. A quick test is to fetch the page without a browser and search for a sentence from it. If the count is zero, the text only exists after JavaScript runs. Our JavaScript SEO guide covers the fixes.
curl -s https://www.example.com/pricing | grep -c "a sentence from your page"What is the status of llms.txt for LLM SEO?
llms.txt is a proposal by Jeremy Howard, published in September 2024: a Markdown file at /llms.txt with a project name, a short summary and links to the pages an LLM should read, meant for use when a model is answering. The honest status is that it is optional and unproven. Google’s guide says Google Search ignores such files and that creating them will neither harm nor help your visibility there. We found no statement from OpenAI, Anthropic or Perplexity that their crawlers read your llms.txt.
That does not make the file useless. It costs little, it is a tidy index of your best pages, and a coding agent that already works on your documentation can use it. Just do not expect it to change citations in ChatGPT or Google. Our llms.txt examples show formats that work, and the free llms.txt generator builds a valid file.
Which content and entity signals help LLM assistants cite you?
No vendor publishes a ranking recipe for assistants, so stay with what is documented or measured.
- Original, specific content. Google says unique, non-commodity content will likely influence your presence in AI search more than any other suggestion in its guide.
- Facts with numbers, quotes and sources. In the GEO research, adding statistics, quotations and source citations raised visibility in the authors’ test setup, while keyword stuffing did not. Use them only where true.
- Clear entities. Use the same name for your company, product and people everywhere. Say what you are in the first paragraph of your home page and About page, and keep structured data consistent with the visible text.
- Answer-first structure. A heading that states the question and a direct answer in the first sentences give a retrieved passage context. See answer engine optimization.
- Genuine mentions. Google says seeking inauthentic mentions across the web is not as helpful as it might seem. Documentation links, honest reviews and independent comparisons are the mentions worth earning.
- Visible dates. OpenAI tells users to check when a cited source was published or updated, so show your update date.
What should an LLM SEO tool or ChatGPT SEO tool do?
Be sceptical of any tool that promises a position inside an assistant, because there is no ranking list to climb. Google also says no third-party tool has access to its internal ranking or AI systems. A useful LLM SEO tool does five things:
- Runs real prompts on named assistants and tells you which ones it covers.
- Stores every answer and its sources, so you can see changes over time.
- Separates a citation (a link to your domain) from a mention (your name in the text).
- Tests whether the assistants’ crawlers can reach your site.
- Never promises guaranteed citations.
Serpel does this for ChatGPT with web search and for Google AI Overviews, records the sources of each answer and tests your robots.txt against the AI crawlers. It does not query Claude, Gemini or Perplexity, so for those assistants you need their own checks. Our comparison of the best AI visibility tools covers the alternatives.
How do you measure LLM SEO?
Build a prompt set
Write 20 to 50 questions your customers ask an assistant, in their words. Include a few comparison and recommendation prompts.
Record a baseline
Run the prompts and note, for each answer, whether you were cited, mentioned or missing, and which domains were cited instead. Repeat the run, because answers vary.
Track the prompts with the Serpel CLI serpel ai add --project <project-id> --prompt "Which SEO tool has a CLI and an MCP server?" serpel ai run --project <project-id> --wait serpel ai status --project <project-id>Read your logs
Count requests from OAI-SearchBot, Claude-SearchBot, PerplexityBot and the user-triggered bots, and check that they receive status 200.
Add the engine reports
Read Search Console’s Generative AI performance report and Bing’s AI Performance report, which count impressions and citations in the products they cover.
Change one thing, then repeat
Edit a page, wait for the crawlers to return and run the same prompts again. Compare against the baseline, not against a single lucky answer.
Frequently asked questions
What is LLM SEO?
LLM SEO is optimising your site so that assistants built on large language models, such as ChatGPT, Claude, Gemini and Perplexity, can find, read and cite it. It covers crawler access, server-rendered content, quotable facts and measurement of citations, and it builds on classic SEO.
Is LLM SEO different from SEO?
Mostly it builds on SEO, and Google says optimising for generative AI search is still SEO. The differences are the extra crawlers to manage, the need for passages that are easy to quote and the way you measure success, which is citations per prompt rather than a ranking position.
Should I block GPTBot and ClaudeBot?
Blocking them opts your content out of future model training and does not remove you from live search. OpenAI says its settings are independent, so you can allow OAI-SearchBot and disallow GPTBot. Whether to opt out of training is a business decision.
Does llms.txt help with LLM SEO?
There is no proof that it does. Google says Google Search ignores llms.txt, and we found no statement from OpenAI, Anthropic or Perplexity that their crawlers read it. The file is cheap to publish and can help coding agents that read your documentation, but do not expect more citations.
How do I know if an LLM cites my website?
Run a fixed set of real prompts on each assistant, record whether the answer links to your domain or only names you, and repeat the checks over time. Add your server logs for AI crawlers and the AI reports in Search Console and Bing Webmaster Tools.
Sources
- OpenAI: Overview of OpenAI crawlers, accessed 10 Oct 2026
- OpenAI Help Center: Searching the web with ChatGPT, accessed 10 Oct 2026
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?, accessed 10 Oct 2026
- Anthropic: Web search tool, accessed 10 Oct 2026
- Perplexity: Perplexity crawlers, accessed 10 Oct 2026
- Google for Developers: Google’s common crawlers, accessed 10 Oct 2026
- Google AI for Developers: Grounding with Google Search, accessed 10 Oct 2026
- Google Search Central: Optimizing your website for generative AI features on Google Search, accessed 10 Oct 2026
- llms.txt: A proposal to standardise on using an /llms.txt file, accessed 10 Oct 2026
- Vercel: The rise of the AI crawler, accessed 10 Oct 2026
- Aggarwal et al., GEO: Generative Engine Optimization (arXiv 2311.09735, KDD 2024), accessed 10 Oct 2026
- Bing Webmaster Guidelines, accessed 10 Oct 2026
- OpenAI: OAI-SearchBot IP ranges (JSON), accessed 10 Oct 2026
- Anthropic: Crawler IP ranges (JSON), accessed 10 Oct 2026
Related reading
- AI crawlers: which to allow, which to block, and how to do itA practical list of AI crawlers: what GPTBot, ClaudeBot and others do, which to allow or block, plus robots.txt examples you can copy and test.
- llms.txt example: 3 complete files for docs, shops and blogsThree complete llms.txt examples for a SaaS docs site, an online shop and a blog, plus the format rules, llms-full.txt and a validation script.
- GEO vs SEO: what is different, what stays the same and how to measure itGEO vs SEO: SEO earns rankings and clicks, GEO earns citations in AI answers. See the differences, what stays the same and how to measure both.
- AI visibilitySerpel is an AI visibility tool: it checks whether ChatGPT and Google AI Overviews cite your website for the questions your customers ask.
