# LLM SEO: how to get found and cited by AI assistants

URL: https://serpel.app/blog/llm-seo

Updated: 2026-10-10

LLM SEO is the practice of making your site easy for assistants built on large language models, such as ChatGPT, Claude, Gemini and Perplexity, to find, read and cite. It works through two paths: models learn from training data collected before release, and assistants with search fetch live pages to ground their answers. Crawler access, server-rendered content and quotable facts decide the second path, and this guide explains both, the honest status of llms.txt and how to measure the result.

## Key takeaways

- LLM SEO has two paths. Training data is fixed when a model is built, so you can only influence future versions. Live retrieval, where an assistant searches or fetches pages at answer time, is the path you can influence now.
- Vendors separate their bots by purpose. OpenAI uses GPTBot for training, OAI-SearchBot for search and ChatGPT-User for user requests. Anthropic uses ClaudeBot, Claude-SearchBot and Claude-User. Perplexity uses PerplexityBot and Perplexity-User. Google-Extended is a robots.txt token, not a crawler.
- Allow the search bots and decide on the training bots separately. User-triggered fetchers may ignore robots.txt, so use each vendor’s published IP ranges to tell real requests from fakes.
- Serve key content in the HTML. Vercel’s 2024 analysis found that major AI crawlers fetched JavaScript files but did not execute them.
- llms.txt is a proposal. Google says Google Search ignores it, and we found no statement from OpenAI, Anthropic or Perplexity that their crawlers read it. Treat it as optional.

## What is LLM SEO?

LLM SEO is search engine optimisation for the assistants that sit on top of large language models (LLMs). When someone asks ChatGPT, Claude, Gemini or Perplexity a question, the answer comes from the model’s training, from pages the assistant retrieves while it answers, or from both. LLM SEO is the work of making sure your site is in the second group, readable by the bots that retrieve it, and worth quoting.

You will also hear it called “SEO for AI”, AI SEO, GEO and AEO. The labels overlap, and Google’s [optimisation guide](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) says optimising for generative AI search is still SEO. Our [GEO vs SEO](https://serpel.app/blog/geo-vs-seo) comparison explains how the terms relate, and the [AI SEO](https://serpel.app/blog/ai-seo) guide covers using AI as a tool for SEO work. This article is about being found and cited by LLM assistants.

## How do LLM assistants find and cite sources?

An LLM on its own is not a search engine. It generates text from patterns it learned in training. To cite a page, the product around the model has to retrieve that page at answer time. The vendors describe this in their documentation.

- **ChatGPT:** OpenAI says ChatGPT [may search the web automatically](https://help.openai.com/en/articles/9237897-searching-the-web-with-chatgpt) when a question benefits from current information, typically rewrites the question into targeted queries and may attach citations. It also warns that results and citations can be incomplete, outdated or incorrect.
- **Claude:** Anthropic’s [web search tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool) lets Claude decide when to search, run the searches and answer with cited sources. Claude answers directly for stable knowledge and searches for current or changing information. The page does not say which index it searches.
- **Gemini:** Google’s [grounding with Google Search](https://ai.google.dev/gemini-api/docs/google-search) lets the model decide whether a search helps, generate queries, run them and answer with citations to the sources.
- **Perplexity:** its [crawler documentation](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) says PerplexityBot is designed to surface and link websites in Perplexity search results, and that Perplexity-User visits a page to answer a user’s question and includes a link.

### Training data versus live retrieval

**The two ways an LLM assistant can know about your site**
| Path | How it works | What you can do | When it takes effect |
| --- | --- | --- | --- |
| Training data | Crawlers collect web content that may be used to train future models. What a released model learned is fixed | Allow or disallow training bots such as GPTBot and ClaudeBot in robots.txt, which is the control the vendors document | Only for future model versions. A block applies to future crawls, not to what a released model already learned |
| Live retrieval | The assistant searches or fetches pages while it answers and cites them | Allow the search bots, stay indexed, serve content as text and write quotable passages | As soon as crawlers revisit. OpenAI says search systems can take about 24 hours to adjust to a robots.txt change |
| User-requested fetch | A user asks about a specific page and the assistant fetches it | Make sure the page loads quickly and returns status 200 to the user-triggered bots | Immediately, on each request |

The practical consequence is that a citation needs retrieval. If an assistant answers from memory it will not link to you, and what it says about your product depends on what was written about it before its training cut-off. That is why most LLM SEO effort goes into retrieval.

## Which crawlers should you allow for LLM SEO?

Each vendor splits its bots by purpose, so you can allow search and block training. This table shows the documented behaviour. Our guide to [AI crawlers](https://serpel.app/blog/ai-crawlers) lists the rest.

**Crawlers of the main LLM assistants and what the vendors say about them**
| Vendor | Bot | Purpose | robots.txt |
| --- | --- | --- | --- |
| OpenAI | OAI-SearchBot | Surfaces websites in ChatGPT’s search features | Applies. Sites that opt out are not shown in ChatGPT search answers |
| OpenAI | GPTBot | Crawls content that may be used to train foundation models | Applies. Disallow it to opt out of training |
| OpenAI | ChatGPT-User | Visits pages for certain user actions | OpenAI says the rules may not apply to user-initiated requests |
| Anthropic | Claude-SearchBot | Improves the quality of Claude’s search results | Honours robots.txt. Disabling it may reduce visibility in search results |
| Anthropic | ClaudeBot | Collects content that could contribute to model training | Honours robots.txt. A block excludes future material from training |
| Anthropic | Claude-User | Retrieves pages when a user asks a question | Honours robots.txt. Disabling it may reduce visibility for user-directed search |
| Perplexity | PerplexityBot | Surfaces and links websites in Perplexity results. Not used to train foundation models | Applies. Perplexity recommends allowing it |
| Perplexity | Perplexity-User | Visits a page to answer a user’s question | Generally ignores robots.txt, because a user requested the fetch |
| Google | Google-Extended | A robots.txt token that controls use of crawled content for Gemini training and grounding | A token only, with no separate user agent. Does not affect inclusion in Search or ranking |

Sources: the documentation of [OpenAI](https://developers.openai.com/api/docs/bots), [Anthropic](https://privacy.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler), [Perplexity](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) and [Google](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers). OpenAI and Perplexity each say their settings work independently, and Anthropic documents its three bots separately, so the file below allows search and user requests and opts out of training.

```text
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /

User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
Allow: /

User-agent: GPTBot
User-agent: ClaudeBot
Disallow: /
```

Google-Extended is left out on purpose. Google says the token covers both Gemini training and grounding, which means serving the Search index to the model at prompt time, so decide on it separately. Whether to opt out of training at all is a business decision, not an SEO one, and it does not decide whether you can be cited in live answers. Allowing the crawler is not enough on its own. A firewall or bot protection can block a bot that robots.txt allows, so check your CDN rules against the vendors’ published IP ranges, such as [OpenAI’s OAI-SearchBot list](https://openai.com/searchbot.json) and [Anthropic’s list](https://claude.com/crawling/bots.json). To see how your file reads, use the free [robots.txt checker](https://serpel.app/tools/robots-txt-checker). It tests your robots.txt against 13 AI crawlers, as every Serpel crawl does.

## Do LLM crawlers run JavaScript?

Mostly not, according to the best public evidence. In a [December 2024 analysis](https://vercel.com/blog/the-rise-of-the-ai-crawler) of its network, Vercel found that none of the major AI crawlers it measured rendered JavaScript, with the exceptions of Gemini, which uses Googlebot’s infrastructure, and Applebot. ChatGPT and Claude fetched JavaScript files (11.50% and 23.84% of their requests) but did not execute them. The data is almost two years old, comes from Vercel’s own network and predates later crawler changes, so read it as a strong hint, not a current guarantee.

The safe rule is to put everything an assistant should quote in the initial HTML. A quick test is to fetch the page without a browser and search for a sentence from it. If the count is zero, the text only exists after JavaScript runs. Our [JavaScript SEO](https://serpel.app/blog/javascript-seo) guide covers the fixes.

```bash
curl -s https://www.example.com/pricing | grep -c "a sentence from your page"
```

## What is the status of llms.txt for LLM SEO?

[llms.txt](https://llmstxt.org/) is a proposal by Jeremy Howard, published in September 2024: a Markdown file at /llms.txt with a project name, a short summary and links to the pages an LLM should read, meant for use when a model is answering. The honest status is that it is optional and unproven. Google’s guide says Google Search ignores such files and that creating them will neither harm nor help your visibility there. We found no statement from OpenAI, Anthropic or Perplexity that their crawlers read your llms.txt.

That does not make the file useless. It costs little, it is a tidy index of your best pages, and a coding agent that already works on your documentation can use it. Just do not expect it to change citations in ChatGPT or Google. Our [llms.txt examples](https://serpel.app/blog/llms-txt-examples) show formats that work, and the free [llms.txt generator](https://serpel.app/tools/llms-txt-generator) builds a valid file.

## Which content and entity signals help LLM assistants cite you?

No vendor publishes a ranking recipe for assistants, so stay with what is documented or measured.

- **Original, specific content.** Google says unique, non-commodity content will likely influence your presence in AI search more than any other suggestion in its guide.
- **Facts with numbers, quotes and sources.** In the [GEO research](https://serpel.app/blog/generative-engine-optimization), adding statistics, quotations and source citations raised visibility in the authors’ test setup, while keyword stuffing did not. Use them only where true.
- **Clear entities.** Use the same name for your company, product and people everywhere. Say what you are in the first paragraph of your home page and About page, and keep structured data consistent with the visible text.
- **Answer-first structure.** A heading that states the question and a direct answer in the first sentences give a retrieved passage context. See [answer engine optimization](https://serpel.app/blog/answer-engine-optimization).
- **Genuine mentions.** Google says seeking inauthentic mentions across the web is not as helpful as it might seem. Documentation links, honest reviews and independent comparisons are the mentions worth earning.
- **Visible dates.** OpenAI tells users to check when a cited source was published or updated, so show your update date.

> **Do not hide instructions for LLMs in your pages:** Hidden text that tries to steer an AI system is a risk, not a shortcut. Microsoft’s Bing Webmaster Guidelines list prompt injection among the practices they treat as abuse. Write for readers and let the assistants quote what they find.

## What should an LLM SEO tool or ChatGPT SEO tool do?

Be sceptical of any tool that promises a position inside an assistant, because there is no ranking list to climb. Google also says no third-party tool has access to its internal ranking or AI systems. A useful LLM SEO tool does five things:

1. Runs real prompts on named assistants and tells you which ones it covers.
2. Stores every answer and its sources, so you can see changes over time.
3. Separates a citation (a link to your domain) from a mention (your name in the text).
4. Tests whether the assistants’ crawlers can reach your site.
5. Never promises guaranteed citations.

[Serpel](https://serpel.app/features/ai-visibility) does this for ChatGPT with web search and for Google AI Overviews, records the sources of each answer and tests your robots.txt against the AI crawlers. It does not query Claude, Gemini or Perplexity, so for those assistants you need their own checks. Our comparison of the [best AI visibility tools](https://serpel.app/blog/best-ai-visibility-tools) covers the alternatives.

## How do you measure LLM SEO?

1. **Build a prompt set** Write 20 to 50 questions your customers ask an assistant, in their words. Include a few comparison and recommendation prompts.

2. **Record a baseline** Run the prompts and note, for each answer, whether you were cited, mentioned or missing, and which domains were cited instead. Repeat the run, because answers vary.

```bash
serpel ai add --project <project-id> --prompt "Which SEO tool has a CLI and an MCP server?"
serpel ai run --project <project-id> --wait
serpel ai status --project <project-id>
```

3. **Read your logs** Count requests from OAI-SearchBot, Claude-SearchBot, PerplexityBot and the user-triggered bots, and check that they receive status 200.

4. **Add the engine reports** Read Search Console’s Generative AI performance report and Bing’s AI Performance report, which count impressions and citations in the products they cover.

5. **Change one thing, then repeat** Edit a page, wait for the crawlers to return and run the same prompts again. Compare against the baseline, not against a single lucky answer.

## Frequently asked questions

### What is LLM SEO?

LLM SEO is optimising your site so that assistants built on large language models, such as ChatGPT, Claude, Gemini and Perplexity, can find, read and cite it. It covers crawler access, server-rendered content, quotable facts and measurement of citations, and it builds on classic SEO.

### Is LLM SEO different from SEO?

Mostly it builds on SEO, and Google says optimising for generative AI search is still SEO. The differences are the extra crawlers to manage, the need for passages that are easy to quote and the way you measure success, which is citations per prompt rather than a ranking position.

### Should I block GPTBot and ClaudeBot?

Blocking them opts your content out of future model training and does not remove you from live search. OpenAI says its settings are independent, so you can allow OAI-SearchBot and disallow GPTBot. Whether to opt out of training is a business decision.

### Does llms.txt help with LLM SEO?

There is no proof that it does. Google says Google Search ignores llms.txt, and we found no statement from OpenAI, Anthropic or Perplexity that their crawlers read it. The file is cheap to publish and can help coding agents that read your documentation, but do not expect more citations.

### How do I know if an LLM cites my website?

Run a fixed set of real prompts on each assistant, record whether the answer links to your domain or only names you, and repeat the checks over time. Add your server logs for AI crawlers and the AI reports in Search Console and Bing Webmaster Tools.

## Sources

- [OpenAI: Overview of OpenAI crawlers](https://developers.openai.com/api/docs/bots), accessed 2026-10-10
- [OpenAI Help Center: Searching the web with ChatGPT](https://help.openai.com/en/articles/9237897-searching-the-web-with-chatgpt), accessed 2026-10-10
- [Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?](https://privacy.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler), accessed 2026-10-10
- [Anthropic: Web search tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool), accessed 2026-10-10
- [Perplexity: Perplexity crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers), accessed 2026-10-10
- [Google for Developers: Google’s common crawlers](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers), accessed 2026-10-10
- [Google AI for Developers: Grounding with Google Search](https://ai.google.dev/gemini-api/docs/google-search), accessed 2026-10-10
- [Google Search Central: Optimizing your website for generative AI features on Google Search](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide), accessed 2026-10-10
- [llms.txt: A proposal to standardise on using an /llms.txt file](https://llmstxt.org/), accessed 2026-10-10
- [Vercel: The rise of the AI crawler](https://vercel.com/blog/the-rise-of-the-ai-crawler), accessed 2026-10-10
- [Aggarwal et al., GEO: Generative Engine Optimization (arXiv 2311.09735, KDD 2024)](https://arxiv.org/abs/2311.09735), accessed 2026-10-10
- [Bing Webmaster Guidelines](https://www.bing.com/webmasters/help/webmaster-guidelines-30fba23a), accessed 2026-10-10
- [OpenAI: OAI-SearchBot IP ranges (JSON)](https://openai.com/searchbot.json), accessed 2026-10-10
- [Anthropic: Crawler IP ranges (JSON)](https://claude.com/crawling/bots.json), accessed 2026-10-10