AI search

LLM SEO: how to get found and cited by AI assistants

LLM SEO is the practice of making your site easy for assistants built on large language models, such as ChatGPT, Claude, Gemini and Perplexity, to find, read and cite. It works through two paths: models learn from training data collected before release, and assistants with search fetch live pages to ground their answers. Crawler access, server-rendered content and quotable facts decide the second path, and this guide explains both, the honest status of llms.txt and how to measure the result.

Serpel Team9 min read

Serpel illustration: a crawler access list for OAI-SearchBot, ClaudeBot, PerplexityBot, GPTBot and Google-Extended, each with its logo and an allowed or blocked status

What is LLM SEO?

LLM SEO is search engine optimisation for the assistants that sit on top of large language models (LLMs). When someone asks ChatGPT, Claude, Gemini or Perplexity a question, the answer comes from the model’s training, from pages the assistant retrieves while it answers, or from both. LLM SEO is the work of making sure your site is in the second group, readable by the bots that retrieve it, and worth quoting.

You will also hear it called “SEO for AI”, AI SEO, GEO and AEO. The labels overlap, and Google’s optimisation guide says optimising for generative AI search is still SEO. Our GEO vs SEO comparison explains how the terms relate, and the AI SEO guide covers using AI as a tool for SEO work. This article is about being found and cited by LLM assistants.

How do LLM assistants find and cite sources?

An LLM on its own is not a search engine. It generates text from patterns it learned in training. To cite a page, the product around the model has to retrieve that page at answer time. The vendors describe this in their documentation.

  • ChatGPT: OpenAI says ChatGPT may search the web automatically when a question benefits from current information, typically rewrites the question into targeted queries and may attach citations. It also warns that results and citations can be incomplete, outdated or incorrect.
  • Claude: Anthropic’s web search tool lets Claude decide when to search, run the searches and answer with cited sources. Claude answers directly for stable knowledge and searches for current or changing information. The page does not say which index it searches.
  • Gemini: Google’s grounding with Google Search lets the model decide whether a search helps, generate queries, run them and answer with citations to the sources.
  • Perplexity: its crawler documentation says PerplexityBot is designed to surface and link websites in Perplexity search results, and that Perplexity-User visits a page to answer a user’s question and includes a link.

Training data versus live retrieval

The two ways an LLM assistant can know about your site
PathHow it worksWhat you can doWhen it takes effect
Training dataCrawlers collect web content that may be used to train future models. What a released model learned is fixedAllow or disallow training bots such as GPTBot and ClaudeBot in robots.txt, which is the control the vendors documentOnly for future model versions. A block applies to future crawls, not to what a released model already learned
Live retrievalThe assistant searches or fetches pages while it answers and cites themAllow the search bots, stay indexed, serve content as text and write quotable passagesAs soon as crawlers revisit. OpenAI says search systems can take about 24 hours to adjust to a robots.txt change
User-requested fetchA user asks about a specific page and the assistant fetches itMake sure the page loads quickly and returns status 200 to the user-triggered botsImmediately, on each request

The practical consequence is that a citation needs retrieval. If an assistant answers from memory it will not link to you, and what it says about your product depends on what was written about it before its training cut-off. That is why most LLM SEO effort goes into retrieval.

Which crawlers should you allow for LLM SEO?

Each vendor splits its bots by purpose, so you can allow search and block training. This table shows the documented behaviour. Our guide to AI crawlers lists the rest.

Crawlers of the main LLM assistants and what the vendors say about them
VendorBotPurposerobots.txt
OpenAIOAI-SearchBotSurfaces websites in ChatGPT’s search featuresApplies. Sites that opt out are not shown in ChatGPT search answers
OpenAIGPTBotCrawls content that may be used to train foundation modelsApplies. Disallow it to opt out of training
OpenAIChatGPT-UserVisits pages for certain user actionsOpenAI says the rules may not apply to user-initiated requests
AnthropicClaude-SearchBotImproves the quality of Claude’s search resultsHonours robots.txt. Disabling it may reduce visibility in search results
AnthropicClaudeBotCollects content that could contribute to model trainingHonours robots.txt. A block excludes future material from training
AnthropicClaude-UserRetrieves pages when a user asks a questionHonours robots.txt. Disabling it may reduce visibility for user-directed search
PerplexityPerplexityBotSurfaces and links websites in Perplexity results. Not used to train foundation modelsApplies. Perplexity recommends allowing it
PerplexityPerplexity-UserVisits a page to answer a user’s questionGenerally ignores robots.txt, because a user requested the fetch
GoogleGoogle-ExtendedA robots.txt token that controls use of crawled content for Gemini training and groundingA token only, with no separate user agent. Does not affect inclusion in Search or ranking

Sources: the documentation of OpenAI, Anthropic, Perplexity and Google. OpenAI and Perplexity each say their settings work independently, and Anthropic documents its three bots separately, so the file below allows search and user requests and opts out of training.

robots.txt that allows AI search and opts out of training
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /

User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
Allow: /

User-agent: GPTBot
User-agent: ClaudeBot
Disallow: /

Google-Extended is left out on purpose. Google says the token covers both Gemini training and grounding, which means serving the Search index to the model at prompt time, so decide on it separately. Whether to opt out of training at all is a business decision, not an SEO one, and it does not decide whether you can be cited in live answers. Allowing the crawler is not enough on its own. A firewall or bot protection can block a bot that robots.txt allows, so check your CDN rules against the vendors’ published IP ranges, such as OpenAI’s OAI-SearchBot list and Anthropic’s list. To see how your file reads, use the free robots.txt checker. It tests your robots.txt against 13 AI crawlers, as every Serpel crawl does.

Do LLM crawlers run JavaScript?

Mostly not, according to the best public evidence. In a December 2024 analysis of its network, Vercel found that none of the major AI crawlers it measured rendered JavaScript, with the exceptions of Gemini, which uses Googlebot’s infrastructure, and Applebot. ChatGPT and Claude fetched JavaScript files (11.50% and 23.84% of their requests) but did not execute them. The data is almost two years old, comes from Vercel’s own network and predates later crawler changes, so read it as a strong hint, not a current guarantee.

The safe rule is to put everything an assistant should quote in the initial HTML. A quick test is to fetch the page without a browser and search for a sentence from it. If the count is zero, the text only exists after JavaScript runs. Our JavaScript SEO guide covers the fixes.

Check that a sentence is in the delivered HTML
curl -s https://www.example.com/pricing | grep -c "a sentence from your page"

What is the status of llms.txt for LLM SEO?

llms.txt is a proposal by Jeremy Howard, published in September 2024: a Markdown file at /llms.txt with a project name, a short summary and links to the pages an LLM should read, meant for use when a model is answering. The honest status is that it is optional and unproven. Google’s guide says Google Search ignores such files and that creating them will neither harm nor help your visibility there. We found no statement from OpenAI, Anthropic or Perplexity that their crawlers read your llms.txt.

That does not make the file useless. It costs little, it is a tidy index of your best pages, and a coding agent that already works on your documentation can use it. Just do not expect it to change citations in ChatGPT or Google. Our llms.txt examples show formats that work, and the free llms.txt generator builds a valid file.

Which content and entity signals help LLM assistants cite you?

No vendor publishes a ranking recipe for assistants, so stay with what is documented or measured.

  • Original, specific content. Google says unique, non-commodity content will likely influence your presence in AI search more than any other suggestion in its guide.
  • Facts with numbers, quotes and sources. In the GEO research, adding statistics, quotations and source citations raised visibility in the authors’ test setup, while keyword stuffing did not. Use them only where true.
  • Clear entities. Use the same name for your company, product and people everywhere. Say what you are in the first paragraph of your home page and About page, and keep structured data consistent with the visible text.
  • Answer-first structure. A heading that states the question and a direct answer in the first sentences give a retrieved passage context. See answer engine optimization.
  • Genuine mentions. Google says seeking inauthentic mentions across the web is not as helpful as it might seem. Documentation links, honest reviews and independent comparisons are the mentions worth earning.
  • Visible dates. OpenAI tells users to check when a cited source was published or updated, so show your update date.

What should an LLM SEO tool or ChatGPT SEO tool do?

Be sceptical of any tool that promises a position inside an assistant, because there is no ranking list to climb. Google also says no third-party tool has access to its internal ranking or AI systems. A useful LLM SEO tool does five things:

  1. Runs real prompts on named assistants and tells you which ones it covers.
  2. Stores every answer and its sources, so you can see changes over time.
  3. Separates a citation (a link to your domain) from a mention (your name in the text).
  4. Tests whether the assistants’ crawlers can reach your site.
  5. Never promises guaranteed citations.

Serpel does this for ChatGPT with web search and for Google AI Overviews, records the sources of each answer and tests your robots.txt against the AI crawlers. It does not query Claude, Gemini or Perplexity, so for those assistants you need their own checks. Our comparison of the best AI visibility tools covers the alternatives.

How do you measure LLM SEO?

  1. Build a prompt set

    Write 20 to 50 questions your customers ask an assistant, in their words. Include a few comparison and recommendation prompts.

  2. Record a baseline

    Run the prompts and note, for each answer, whether you were cited, mentioned or missing, and which domains were cited instead. Repeat the run, because answers vary.

    Track the prompts with the Serpel CLI
    serpel ai add --project <project-id> --prompt "Which SEO tool has a CLI and an MCP server?"
    serpel ai run --project <project-id> --wait
    serpel ai status --project <project-id>
  3. Read your logs

    Count requests from OAI-SearchBot, Claude-SearchBot, PerplexityBot and the user-triggered bots, and check that they receive status 200.

  4. Add the engine reports

    Read Search Console’s Generative AI performance report and Bing’s AI Performance report, which count impressions and citations in the products they cover.

  5. Change one thing, then repeat

    Edit a page, wait for the crawlers to return and run the same prompts again. Compare against the baseline, not against a single lucky answer.

Frequently asked questions

What is LLM SEO?

LLM SEO is optimising your site so that assistants built on large language models, such as ChatGPT, Claude, Gemini and Perplexity, can find, read and cite it. It covers crawler access, server-rendered content, quotable facts and measurement of citations, and it builds on classic SEO.

Is LLM SEO different from SEO?

Mostly it builds on SEO, and Google says optimising for generative AI search is still SEO. The differences are the extra crawlers to manage, the need for passages that are easy to quote and the way you measure success, which is citations per prompt rather than a ranking position.

Should I block GPTBot and ClaudeBot?

Blocking them opts your content out of future model training and does not remove you from live search. OpenAI says its settings are independent, so you can allow OAI-SearchBot and disallow GPTBot. Whether to opt out of training is a business decision.

Does llms.txt help with LLM SEO?

There is no proof that it does. Google says Google Search ignores llms.txt, and we found no statement from OpenAI, Anthropic or Perplexity that their crawlers read it. The file is cheap to publish and can help coding agents that read your documentation, but do not expect more citations.

How do I know if an LLM cites my website?

Run a fixed set of real prompts on each assistant, record whether the answer links to your domain or only names you, and repeat the checks over time. Add your server logs for AI crawlers and the AI reports in Search Console and Bing Webmaster Tools.

Sources

  1. OpenAI: Overview of OpenAI crawlers, accessed 10 Oct 2026
  2. OpenAI Help Center: Searching the web with ChatGPT, accessed 10 Oct 2026
  3. Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?, accessed 10 Oct 2026
  4. Anthropic: Web search tool, accessed 10 Oct 2026
  5. Perplexity: Perplexity crawlers, accessed 10 Oct 2026
  6. Google for Developers: Google’s common crawlers, accessed 10 Oct 2026
  7. Google AI for Developers: Grounding with Google Search, accessed 10 Oct 2026
  8. Google Search Central: Optimizing your website for generative AI features on Google Search, accessed 10 Oct 2026
  9. llms.txt: A proposal to standardise on using an /llms.txt file, accessed 10 Oct 2026
  10. Vercel: The rise of the AI crawler, accessed 10 Oct 2026
  11. Aggarwal et al., GEO: Generative Engine Optimization (arXiv 2311.09735, KDD 2024), accessed 10 Oct 2026
  12. Bing Webmaster Guidelines, accessed 10 Oct 2026
  13. OpenAI: OAI-SearchBot IP ranges (JSON), accessed 10 Oct 2026
  14. Anthropic: Crawler IP ranges (JSON), accessed 10 Oct 2026

Related reading