AI search ranking factors: what is documented and what is speculation
The engines themselves document only a short list of AI search factors, and nearly all of it is about eligibility: can the right crawler reach the page, is the page in the index the engine searches, and have you allowed its text to be shown. Which of the eligible pages gets cited is not published by Google, OpenAI, Perplexity, Microsoft or Anthropic. Everything said about selection is inference from testing, or speculation.
"Ranking factor" is itself a loose term here. AI answers cite a handful of sources without a visible order, and the set changes between runs. Below, each common claim carries one of three labels. Documented means an engine's own documentation says it. Inferred means it follows from how the systems are known to work, or from repeated independent testing, but no engine confirms it. Speculative means there is no good evidence either way.
Documented by the engines
| Factor | Label | Source |
|---|---|---|
| Page is indexed and eligible to show a snippet in Google Search | Documented | Google, "AI features and your website" |
nosnippet, data-nosnippet and max-snippet limit what AI Overviews and AI Mode can show | Documented | Google, same page |
| Google-Extended does not affect Search or its AI features, but covers Gemini training and Gemini Apps grounding | Documented | Google's common crawlers documentation |
| OAI-SearchBot must be allowed for pages to appear in ChatGPT search answers | Documented | OpenAI crawler overview and publisher FAQ |
| PerplexityBot must be allowed to be surfaced in Perplexity search | Documented | Perplexity bot documentation |
| Blocking Claude-SearchBot may reduce visibility in Claude's search results | Documented | Anthropic's crawler help article |
| Copilot answers are grounded in Bing, and IndexNow helps Bing see updates quickly | Documented | Microsoft Bing Webmaster guidance |
Bing's noarchive keeps a page out of its chat answers; nocache limits what is shown | Documented | Bing Webmaster blog, 2023 |
Notice what these have in common: they tell you how to be allowed in, never how to be picked. That is the entire documented list for most engines. The crawler pages on RankWave AI give each agent's token and purpose.
Inferred from how the systems work
- Ranking well in the underlying index helps. Inferred. Every major engine retrieves by searching first, so a page that ranks for the question, or for one of the sub-questions an engine runs behind the scenes, is more likely to be read. Google describes this query fan-out for AI Mode and AI Overviews.
- Citations overlap with Bing for ChatGPT and with Brave for Claude. Inferred. OpenAI says it uses third-party search providers; Anthropic lists Brave Search as a provider. Independent testers have reported overlapping citations, but neither company publishes the weighting.
- Direct, self-contained passages get quoted more. Inferred. A retrieval system lifts passages. A passage that answers without depending on the rest of the page is easier to use. This is consistent with the 2023 GEO research and with Google's advice to put important content in text.
- Freshness matters for time-sensitive questions. Inferred. Engines ground answers to get current facts; an outdated page gives them a reason to look elsewhere.
- Brand recognition carries weight. Inferred. Models blend retrieved pages with what they learned in training, and sites that are widely cited elsewhere tend to be cited again. How much this weighs is unknown.
Speculative, or contradicted
- Schema markup directly raises your chance of an AI citation. Speculative. Structured data helps search engines understand a page, and Google recommends it match visible content, but Google also says no special markup is needed for its AI features. No LLM operator documents using it for selection.
- llms.txt improves AI visibility. Speculative. It is a proposal. No major AI search provider has said it reads the file when choosing sources. See the llms.txt glossary entry.
- Ideal word counts or passage lengths. Speculative. Figures like "answers of exactly 40 to 60 words" come from correlation studies, not engine documentation.
- Allowing training crawlers earns search citations. Contradicted for OpenAI and Anthropic, whose search and training agents are controlled separately.
- Buying placement. No engine offers paid entry into organic answers. OpenAI has said its ads do not influence answers.
How to use this list
Act on the documented factors first, because they are cheap and certain: a crawler audit, an indexing check, and a scan for snippet directives you did not mean to set. Hydrogen SEO outputs its titles, meta tags and schema in server-rendered HTML, and its robots.txt editor lets you set per-agent rules from wp-admin, which covers most of that list on a WordPress site.
Then pick one inferred factor at a time and test it on a handful of pages, using the experiment design on the GEO strategy page. Leave speculative tactics alone unless they cost almost nothing, and never let one replace work on the first two tiers. When someone shows you an "AI ranking factors study", ask which of these three labels its claims deserve.
Common questions
Is there an official list of AI search ranking factors?
No. Engines document how to be eligible, such as which crawler to allow and which snippet rules apply, but none publishes how it chooses which eligible pages to cite.
Does schema markup help ChatGPT cite me?
OpenAI has not said it uses schema for citations. Accurate schema still helps search engines understand your pages, and ChatGPT search draws on search indexes, so it is worth keeping correct.
Do backlinks matter for AI answers?
Indirectly, as far as anyone can tell. Links help pages rank in the indexes that AI engines search, and widely referenced brands tend to be cited more, but no engine documents links as a citation factor.