Generative engine optimization: a strategy, and what the evidence really shows
A sensible GEO strategy is ordinary SEO plus content that hands a model something specific to quote: figures with a source, first-hand results, named methods and clear statements with their conditions attached. That is about as far as the evidence goes. There is one widely cited academic study, run mostly on a research setup rather than the live products, a stream of vendor reports that show correlation at best, and first-party documentation that explains eligibility but not selection.
So the strategy has two halves. Do the things the engines document, because those are settled. Treat everything else as a hypothesis, and test it on your own pages with a method you would trust if someone else reported it.
What the research actually tested
The term comes from "GEO: Generative Engine Optimization", a paper by researchers at Princeton, Georgia Tech, the Allen Institute for AI and IIT Delhi, posted in late 2023 and presented at the KDD conference in 2024. The authors built a benchmark of queries and a generative engine for the experiment, then edited source pages in set ways and measured how visible each page became in the generated answers.
Edits that added quotations, statistics and references to sources did best in their setup, while keyword stuffing did poorly. Results also varied by subject area. That is a useful signal, but read it with its limits in mind: the engine was a research construction using models from 2023, "visibility" was a metric the authors defined, and today's production systems search, rank and cite in ways their operators do not publish. The paper supports a direction (specific, attributable content helps), not a recipe.
How to read vendor studies
Most "AI ranking factor" reports come from tool vendors who collect large samples of AI answers and look for traits the cited pages share. Before acting on one, ask four questions:
- Is the method published? Prompt list, engines, dates, locations, and whether answers were collected logged in or out.
- Is it correlation? Cited pages being long, or old, or linked from many sites does not mean those traits caused the citation. Strong sites tend to have all of them.
- Is the sample relevant? A study of software queries says little about local services or recipes.
- Does the vendor sell the fix? Not disqualifying, but a reason to look harder.
The sources worth weighting most are the engines' own pages: Google's documentation on AI features and its May 2025 Search Central post asking site owners to focus on unique, non-commodity content, OpenAI's crawler and publisher documentation, and the bot pages from Anthropic, Perplexity and Microsoft. They are narrow, but they are true.
A strategy in three layers
Layer 1, settled. Be crawlable by each engine's search agent, indexed in Google and Bing, and eligible to show a snippet. Keep main content in server-rendered HTML. Every engine's documentation supports some version of this.
Layer 2, consistent with the evidence. Write sections that make one claim each and back it with a number, a date, a source or a first-hand observation. Publish information that only you have: test results, pricing you set, data from your own customers, photos of your own work. Both the research and Google's guidance point the same way here.
Layer 3, experiments. llms.txt, schema written for language models, exact passage lengths, statistics density, FAQ blocks on every page. None of these is documented to affect selection by any major engine. Some are cheap, which is fine, but give them a test before you roll them out site-wide. The ranking factors page labels the common claims one by one.
Run your own GEO test
A small, honest experiment beats a large dubious study. Here is a design a single site can run:
- Pick 20 prompts where AI answers already cite someone in your space.
- Record a baseline. Run each prompt three times in each engine you care about, in a clean session, and log the cited domains.
- Split your pages. Choose ten pages that answer those prompts. Change five in one specific way, such as adding sourced figures and a comparison table, and leave five similar pages untouched.
- Wait for a recrawl. Server logs show when OAI-SearchBot, PerplexityBot, Bingbot or Googlebot fetched the edited URLs. Start the clock from there.
- Re-run the prompts at four and eight weeks, same wording, same sessions.
- Compare the two groups, not single pages. If edited and untouched pages move together, the change probably was not the cause.
Expect noise. Answers vary between runs, and engines update their systems without notice during your test window. A result that holds across two cycles is worth rolling out; a one-off jump is not.
Claims to treat with suspicion
Be wary of a "GEO score" that promises to predict citations, of guarantees of placement in ChatGPT or AI Overviews, of fixed rules like "AI prefers 2,000-word pages", and of services selling mentions on sites built only to be scraped. Engines have every incentive to discount content made to game them, and Google's spam policies on scaled content abuse already cover mass-produced pages. If a tactic only works while nobody notices it, it is not a strategy.
Common questions
Is GEO backed by research?
Partly. One academic paper found that adding quotations, statistics and cited sources improved visibility in its test setup. It was not a study of today's production engines, so treat it as a direction rather than proof.
Should I add statistics to every page?
Only where you have real, sourced figures that answer the question. Invented or padded numbers damage trust and can be quoted back at you incorrectly.
How long does a GEO change take to show up?
Allow several weeks. The engine has to recrawl the page, and answers vary from run to run, so you need repeated checks before you can see a change.
Is GEO different from AEO?
In practice they overlap almost entirely. GEO is used more for long answers written by generative models, AEO for direct answers of any kind.