seojuice

LLM SEO: How to Show Up in AI Answers

Vadim Kravcenko
Vadim Kravcenko
Jul 19, 2026 · 9 min read

TL;DR: LLM SEO is the work of becoming a source inside answers from ChatGPT, Claude, Gemini, and Perplexity, rather than winning another blue-link position. Start with retrieval: keep the correct AI search bots unblocked, remain indexable, publish passages that are easy to extract and cite, build credible evidence beyond your own domain, and test whether assistants actually mention you. There is no confirmed public formula for ranking first in an AI answer.

The central mistake is treating an LLM like Google with a chat box attached.

A search engine usually returns ranked documents. An LLM synthesizes an answer from what it learned during training, what it retrieves from the live web, and whatever index supports that retrieval. Your page might influence the answer without receiving a click. It might also rank well on Google and never appear in the generated response.

There is no confirmed public ranking-factor list for ChatGPT, Claude, Gemini, or Perplexity. Anyone promising a repeatable “rank number one in ChatGPT” formula is selling certainty the vendors have not provided (I would like that dashboard too).

Question Classic Google SEO LLM SEO
What are you trying to win? A position among search results A mention or citation inside a generated answer
What does the user receive? A list of links One synthesized response, sometimes with citations
What remains essential? Indexability, useful pages, authority, links Indexability, crawlability, useful pages, authority, references
What changes? The page and its ranking are the primary units A passage, fact, or brand can become the surfaced unit
How do you measure it? Rankings, impressions, clicks Prompt tests, mentions, citations, source inclusion
How transparent is it? Established documentation and reporting tools No vendor-published citation ranking factors
The three ways an LLM can surface your brand: pre-training, live retrieval, and the web index behind the tools.

What LLM SEO actually means

LLM SEO, also called LLM optimization, LLMO, GEO, AISO, or answer-engine optimization, is the practice of making a company and its content more likely to surface inside an AI-generated answer.

The academic term is generative engine optimization. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, and Deshpande described GEO as “the first novel paradigm to aid content creators in improving their content visibility in generative engine responses.” Their Princeton-led paper was accepted to KDD 2024.

The naming is unsettled. I use “LLM SEO” here for work aimed specifically at assistants such as ChatGPT, Claude, Gemini, and Perplexity. AI search optimization is broader and includes answer surfaces such as Google AI Overviews and AI Mode.

The economic outcome is different from a normal search result. Pew Research Center analyzed 68,879 Google searches made by 900 U.S. adults in March 2025. About 18% produced an AI summary. Users clicked a traditional result in 8% of visits containing a summary, compared with 15% when no summary appeared. A link inside the summary received a click in only 1% of visits.

Those figures concern Google summaries, not ChatGPT or Claude. They still demonstrate the structural change: the generated answer can replace the click. Being named in that answer may carry value even when referral traffic is weak or impossible to attribute.

The three places an LLM gets its answer

You cannot make sensible LLM optimization decisions until you separate training from retrieval. They involve different systems, crawlers, and timelines.

1. Pre-training data: frozen memory

Foundation models learn patterns and information from large text collections during training. OpenAI says GPTBot crawls content that may be used to train its generative AI foundation models. Anthropic identifies ClaudeBot as a crawler used for collecting web content for model training.

You do not rank inside this layer. Your company may have been represented in the training corpus, represented poorly, or absent. Publishing a correction this afternoon does not rewrite a completed training run.

Broad, consistent documentation across the web may increase the chance that a future model encounters your brand. The exact inclusion and weighting mechanics are not public, however. I trust the distinction between training and retrieval; I do not trust claims that twelve mentions, twenty links, or one Wikipedia entry will produce a predictable change in model memory.

2. Live retrieval: the layer you can affect fastest

When an assistant searches or browses, it can fetch current documents and use them to ground an answer. This is commonly called retrieval-augmented generation, or RAG. Citations normally come from this retrieval process, not from the model revealing individual documents inside its training corpus.

This is where most practical LLM SEO work belongs. A page must be available to the relevant search crawler, reachable at answer time, and suitable for extraction. If the system cannot retrieve a document, it cannot use that document as a live citation.

OpenAI assigns search discovery to OAI-SearchBot, while ChatGPT-User handles user-triggered visits. Perplexity distinguishes PerplexityBot, which surfaces and links websites in its results, from Perplexity-User, which fetches pages in response to questions. Anthropic similarly lists Claude-SearchBot and Claude-User.

3. The underlying web index

Retrieval needs a collection of documents to search. That may be an underlying index from Google or Bing, the assistant vendor’s own index, or a changing mixture of sources. Provider relationships and retrieval systems move quickly, so do not build permanent infrastructure assumptions around one search partnership.

The durable conclusion is simpler: normal indexability still matters. If search engines cannot discover, render, canonicalize, or index a page, systems drawing from those indexes have nothing useful to retrieve.

Google makes this explicit for its own AI features:

“There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.”

Google Search Central also says that sites do not need new machine-readable files, AI text files, or special markup to appear in those features. That statement applies to Google AI Overviews and AI Mode, not to ChatGPT, Claude, or Perplexity. It still punctures a lot of unnecessary product theatre.

Classic SEO has not become obsolete. It became infrastructure (or, more precisely, part of the infrastructure).

Allow the bots that serve answers

A blanket “block AI” rule is blunt. Training crawlers, search crawlers, and user-triggered fetchers perform different jobs. You can opt out of training without necessarily removing yourself from live answer retrieval.

Vendor Bot Function Vendor’s robots.txt position
OpenAI GPTBot Potential model-training content Respects opt-out directives
OpenAI OAI-SearchBot Surfaces sites in ChatGPT search Respects robots.txt
OpenAI ChatGPT-User User-triggered page visit Rules may not apply
Perplexity PerplexityBot Indexes pages for search results, not foundation-model training Respects robots.txt
Perplexity Perplexity-User User-triggered page fetch Generally ignores robots.txt
Anthropic ClaudeBot Training-content collection Honors robots.txt
Anthropic Claude-SearchBot Search-result retrieval Honors robots.txt
Anthropic Claude-User User-triggered access Honors robots.txt

OpenAI’s documentation is explicit about blocking OAI-SearchBot: opted-out sites “will not be shown in ChatGPT search answers,” although they can still appear as navigational links. Perplexity says PerplexityBot is designed to surface and link websites in its results and is not used to crawl content for foundation-model training.

Make three separate decisions:

  • Training policy: Decide whether GPTBot and ClaudeBot may collect content for possible model training.
  • Search visibility: Allow OAI-SearchBot, PerplexityBot, and Claude-SearchBot if you want retrieval visibility.
  • User fetches: Ensure the site remains technically reachable when a user explicitly asks an assistant to access it.

Blocking a training crawler can be a defensible publishing decision. Blocking search crawlers while expecting citations is contradictory. User-triggered fetches complicate enforcement because OpenAI says robots.txt may not apply to ChatGPT-User, while Perplexity says Perplexity-User generally ignores it.

Review these rules periodically. User-agent names and vendor policies can change.

Make the page easy to quote, not merely easy to rank

Once retrieval works, the next job is giving the model a useful passage to lift.

The strongest peer-reviewed experimental evidence comes from the KDD 2024 GEO paper. Its GEO-bench included 10,000 queries, and the researchers tested content interventions against simulated generative engines. They reported that GEO could boost visibility by up to 40% on their benchmark.

That is an “up to” benchmark result, not a promise of 40% more citations in production ChatGPT. According to the paper’s Table 1, quotation addition produced a 27.8% improvement and statistics addition produced a 25.9% improvement on its position-adjusted word-count metric. Quotations, statistics, and cited sources were among the strongest methods tested.

The same paper found keyword stuffing had “little to no performance improvement” in generative-engine responses. Repeating “best invoicing software” seventeen times is not an LLM strategy. It was barely a credible search strategy before.

A page built for extraction usually includes:

  • A direct answer immediately below the relevant heading.
  • Specific claims attributed to named, credible sources.
  • Tables comparing concrete attributes without hiding the conclusion.
  • Definitions that still make sense outside the surrounding article.
  • Statistics with dates, populations, and enough context to prevent misleading reuse.
  • Stable URLs and crawlable text instead of essential facts trapped inside interfaces.

I would prioritize factual density over raw length. A long article can still be hard to cite if every answer is buried under throat-clearing. We have made that mistake ourselves, adding words because a content brief expected them and leaving the useful sentence six paragraphs below the heading (not our finest editorial instinct).

Schema can help machines interpret page entities and content types, but do not confuse it with a secret AI-citation switch. Google explicitly says no special markup is required for its AI features. Structure should improve clarity for readers and crawlers first.

Build evidence outside your own domain

Your own website is necessary, but it is not the entire evidence set an assistant can retrieve.

In Aleyda Solis’s analysis of SaaS AI-search citations, third-party sources generated 84% to 93% of citation weight across the subverticals and platforms she studied. Her conclusion is blunt:

“Most AI citation weight comes from outside your own website.”

This is a measured pattern in SaaS categories, not a vendor-confirmed ranking factor. It tells us where citations appeared in her analysis; it does not prove that publishing one community mention causes an assistant to recommend a brand.

The practical implication still matters. Documentation, established publications, reviews, comparisons, community discussions, videos, and peer-software sites create independent evidence. An assistant has more material to work with when several credible sources agree on what a company does than when every claim originates on that company’s homepage.

Our guide to where LLMs find your brand examines those external surfaces in more detail.

Do not translate this into synthetic mention spam. Better off-site work looks like this:

  1. Publish original data, documentation, examples, or definitions worth referencing.
  2. Earn inclusion in relevant comparisons rather than manufacturing dozens of thin listings.
  3. Keep product facts consistent across your website and third-party profiles.
  4. Give writers one stable source URL for important claims.
  5. Correct material inaccuracies where you have a legitimate route to do so.

SEOJuice is a two-person team, Lida and me. We cannot outspend large SaaS companies on PR, so our more realistic lever is being referenceable: clear product documentation, reproducible examples, and URLs that survive operational changes.

That last point became concrete when we migrated seojuice.io to seojuice.com in January 2026. A domain move is not just a Google rankings project. Redirects, canonicals, internal links, and updated third-party references also preserve the source trail assistants may retrieve (and yes, migration cleanup took longer than the neat checklist implied).

A practical LLM optimization order of operations

Work in dependency order. Content polishing cannot rescue a page excluded from retrieval.

Priority Action Why it comes now
1 Audit robots.txt and bot access Search-crawler access is a hard retrieval gate
2 Confirm Google, Bing, and site indexability Retrieval may depend on an underlying web index
3 Select prompts tied to customer decisions You need a defined answer surface to measure
4 Rewrite weak pages into direct, sourced passages Clear passages are easier to extract and attribute
5 Strengthen third-party references External corroboration appears useful, but is not a confirmed factor
6 Retest across assistants Visibility varies by model, mode, and prompt wording

Choose prompts that represent an actual decision: “Which tools automatically add internal links to a SaaS blog?” is more diagnostic than typing your brand name and asking the model to review it. Branded prompts mostly test whether the assistant can find you; category and comparison prompts test whether it selects you.

Record the assistant, model or mode, exact prompt, date, browsing status, brands mentioned, cited URLs, and the order of recommendations. A model answering from training memory is not the same test as an assistant retrieving live sources.

Expect noise. Generated answers change, personalization may interfere, and there is no equivalent of Google Search Console reporting every AI impression. I have mistaken a handful of successful prompts for a trend before. Screenshots look persuasive; five observations are still five observations.

Where SEOJuice fits

SEOJuice cannot guarantee that an LLM will cite you. No external tool controls the assistant’s retrieval and synthesis systems.

What we automate is part of the prerequisite layer. SEOJuice works on a live site and continuously applies internal links, meta titles and descriptions, schema markup, and image alt text. Those improvements support discoverability and clean site structure, but they should not be sold as secret ChatGPT ranking factors.

For measurement, our AI visibility checker tests whether AI engines mention or cite your brand for selected prompts. It shows where you appear and where you are absent; it does not manufacture a citation.

If that is the gap you need to diagnose, see if ChatGPT and Perplexity mention your brand with the free AI Visibility Checker. SEOJuice has a free plan and requires no credit card.

The broader shift is not that SEO has died. Classic SEO is necessary but insufficient because retrieval still depends on the open web, while the desired output has moved from ranking a link to becoming part of the answer. That is the practical transition from SEO to GEO.

Frequently Asked Questions

What is LLM SEO?

LLM SEO is the practice of optimizing a brand and its content so assistants such as ChatGPT, Claude, Gemini, and Perplexity can surface and cite them inside generated answers. It is also called LLMO, GEO, AISO, or answer-engine optimization. The primary outcome is a mention or citation rather than a blue-link position.

How do LLMs decide what to cite?

An answer can combine pre-training data, live web retrieval, and an underlying index such as Google, Bing, or the vendor’s own. For current citations, concentrate on live retrieval: allow the relevant search bots, remain indexable, and publish clear, sourced passages. Beyond those access gates, vendors have not published their citation-selection formulas.

How is LLM SEO different from Google SEO?

Classic SEO usually targets a ranked search result. LLM SEO targets inclusion inside one synthesized answer. Indexability, authority, and links still support retrieval, but keyword-density tactics do not transfer. Measurement also shifts from positions and clicks toward repeated prompt tests, mentions, cited sources, and recommendation inclusion.

How do I let ChatGPT and Perplexity find my site?

Do not block their search crawlers: OAI-SearchBot for ChatGPT and PerplexityBot for Perplexity. OpenAI and Perplexity separately identify ChatGPT-User and Perplexity-User as user-triggered fetchers. Training crawlers such as GPTBot are separate from search retrieval, so you can make different access decisions for training and search.

Should I block AI bots from my site?

Separate the training decision from the visibility decision. Blocking GPTBot or ClaudeBot opts content out of those vendors’ training collection. Blocking OAI-SearchBot, PerplexityBot, or Claude-SearchBot can remove the site from their retrieval surfaces. A site seeking AI citations should generally keep relevant search bots accessible while making its own policy decision about training crawlers.

Can I guarantee that an LLM will cite my company?

No. You can control crawlability, indexability, factual clarity, and the quality of your external evidence. Experimental work suggests that quotations, statistics, and cited sources can improve generative-answer visibility, while practitioner analysis shows substantial citation weight on third-party sites. Neither provides a guaranteed citation, and no vendor has published a complete ranking-factor list.