ktau← Research
Research

The AEO playbook doesn't survive a fair test. The model's memory does.

Open access · data

October 4, 2026

12

min read

ktau research team

Empirical

We followed 158,099 AI answers and 468,045 citations for 90 days across ChatGPT, Gemini and Perplexity, then replayed every question through the engines' own APIs. FAQ blocks, schema markup, summary boxes and author bylines stop mattering the moment you compare pages of the same website. What decides a citation is whether the engine's own search retrieves you — and whether the model already knows your brand.

Summary

Generative search engines now decide which web pages a buyer sees as sources. A growing industry tells publishers how to earn those citations: add an FAQ block, add structured data, add a summary box, add an author, make the page faster. Almost none of that advice has been tested outside a controlled setting, and the observational comparisons that circulate are easy to mislead — the practices are concentrated on websites that are cited for reasons that have nothing to do with the practice.

So we tested it properly. Over 90 days we collected every answer the three engines gave to 651 commercial questions in a US consumer market (home internet) and 55 questions in a niche B2B software market, resolved every cited URL to its page and its website, downloaded the pages, and then replayed every question through the engines' own APIs to see the searches they ran, the pages they ranked, and which of those they cited. Finally, we ran controlled experiments on the models behind the engines with the sources held fixed and one factor changed at a time.

Six findings:

  1. Across websites, every recommended practice appears to work. Within a website, none of them does. An FAQ block goes with 18% more citations across all pages — and 3% fewer once you compare pages of the same website. Schema markup: +7.5% across, −5.4% within. Author byline: +14.9% across, −4.6% within. Page speed: +11.7% across, −1.1% within. None of the within-website effects is significant.
  2. Content matters in exactly one way: pages whose title and paragraphs answer the questions people ask are cited more. A one-standard-deviation better paragraph match is worth +15.6% citations; a better title +14.1%. On third-party pages the effects are larger still (+25.7% and +23.8%). These are the only page features that pass all five of our robustness tests.
  3. Page properties explain almost nothing. Retrieval explains ten times more. Everything we could measure about a page accounts for 2.8% of the variation in ChatGPT's citations within a website. The engine's own search accounts for 27%. Every one of the 8,646 URLs ChatGPT cited was a page its own search had retrieved — and it cites 49% of the pages it ranks first against 3% of those at ranks 11–20.
  4. The engines are not Google. ChatGPT rewrites a question into about four keyword-style searches, adds the current year to a quarter of them and a brand the user never named to another quarter. 97% of the pages it retrieves are absent from Google's top 50 for the same question, and 99% from Bing's top 30. Even Gemini, grounded on Google Search, cites only 5.6% of Google's top 10. A classic search rank is not a proxy for a citation.
  5. The model's memory is in the loop before and after the search. When ChatGPT already knows a brand, it writes that brand into its own searches 60% of the time (7% when it doesn't). Websites the model would name without searching are ranked 1.40× higher and have 1.58× the odds of being cited at equal rank; 47% of ChatGPT's citations come from such sites, against a 25% baseline. Nearly half (46%) of the brands ChatGPT names are not supported by any cited page at all, and a brand in the model's prior is named with 5.8× the odds at equal citation support. In controlled experiments, revealing the brand names behind identical product facts reorders the model's ranking 2.1× (ChatGPT) to 3.3× (Gemini) more than simply re-running the prompt — and a well-known brand can lose up to 33 points of top-three rate the moment its name is attached to a profile the model had ranked first anonymously.
  6. Citations drift. Two-week windows of the same question share 63–67% of their within-website variation; ten weeks apart they share 14% (ChatGPT), 29% (Gemini) and 0% (Perplexity). A citation measured once is a snapshot, not a standing.

The niche B2B market replicates all of it, with one difference that matters: there, the models barely know the vendors, website authority has no effect at all, and a single third-party page that names the brand raises the odds of being named by 40× in ChatGPT and 18× in Gemini.

How we measured

Two markets, one method. The main dataset is a US consumer market dominated by a few large brands: 651 questions across customer support, equipment, competitive landscape, speed, pricing, performance and reliability, asked daily from 24 June to 21 September 2026 — 143,249 answers and 396,196 citations from ChatGPT (GPT-5.4 nano with web search), Gemini (Gemini 3.1 Flash Lite with Google Search grounding) and Perplexity (Sonar). The second dataset is a niche B2B market of small vendors of proposal, bid and business-development software for architecture, engineering and construction firms: 55 questions, 14,850 answers, 71,849 citations.

Within-website comparison. A page on a popular website is cited more than a page on an obscure one for reasons that have nothing to do with the page. So every effect is estimated by comparing pages of the same website with each other, with a separate intercept for each of 1,117 websites. That single design decision is what makes the recommended practices disappear.

Five tests, not one. A feature counts as real only if it passes five independent checks: it survives false-discovery control over all fourteen features; it is selected in at least 60% of 200 bootstrap Lasso fits; it holds in a double machine-learning estimate with gradient-boosted trees controlling for every other feature; its partial-effect curve is monotone or single-turn; and it keeps its sign in at least four of five variants of the outcome. Only three features pass: title match, best-paragraph match and word overlap with the questions.

Replay. We sent the 651 questions back through the engines' APIs five times over three days with search enabled, and recorded the search queries each engine wrote, the ranked list of pages it retrieved, and which it cited. For ChatGPT, 100% of the URLs cited in the collection were pages the replay retrieved. The replay also lets us observe the model's prior directly: we ask each engine's model, without search, which brands and websites it would recommend for every question, and compare those to what it later searches for and cites.

Controlled experiments. On the models behind the engines (without search), we fixed a set of sources and varied one factor at a time: the position of a source in the list the model reads, whether the brand and publisher names are shown or masked, and whether a product is added or removed from the set — 60 packets of five competing products described by identical fixed profiles, 1,920 answers.

Finding 1 — Across websites, every trick works. Within a website, none does.

  • Has FAQ block: +18.1% across all pages → −3.4% within the same website
  • Has summary box: +20.7% across all pages → +6.7% within the same website
  • Has author byline: +14.9% across all pages → −4.6% within the same website
  • Has schema.org markup: +7.5% across all pages → −5.4% within the same website
  • Faster page (per +1 SD): +11.7% across all pages → −1.1% within the same website

Percent more citations for a page with the practice versus one without. None of the within-website estimates is statistically significant.

The explanation is Simpson's paradox. FAQ blocks appear on 43% of the pages of the most-cited quarter of websites and on 23% of the pages of the least-cited quarter. Compare across websites and you credit the FAQ block with the website's reputation, size and topical focus. Compare within websites and the credit evaporates. The same pattern holds on third-party pages alone, in the second market, and for every engine separately.

This does not mean the practices are useless — an FAQ block may serve readers or classic search perfectly well. It means our data give no reason to expect citations from AI engines because of them.

Finding 2 — Pages are cited for the questions they answer

The three features that survive all five tests measure one thing: how well the page's text matches the questions people ask. One standard deviation better in the paragraph that best matches the questions earns 15.6% more citations (95% CI 11.2–20.2%); in the title, 14.1%; in plain word overlap, 9.1%. On third-party pages — the ones most publishers can actually influence — the effects grow to 25.7%, 23.8% and 12.9%.

Matching also decides which questions a page is cited for. For every page and every query, the query the page is cited for matches it better than one it is not cited for in 88% of comparisons (within-page AUC 0.878; ChatGPT 0.873, Gemini 0.880, Perplexity 0.894). Length, outbound links, headings, lists and tables have no effect at all; there is no page length that works best.

Finding 3 — Retrieval decides; the page explains almost nothing

Of the variation in citation counts between pages of the same website, 72% is stable and repeatable — the same pages win in odd and even days, with a reliability of 0.90 over 90 days. Page features explain 2.8% of it for ChatGPT. The engine's own retrieval explains 27%.

The replay shows why. For each question ChatGPT consults about 26 pages and cites 3. It cites 48.7% of the pages it ranks first, 36.8% of those ranked second, 12.3% of ranks 6–10 and 3.3% of ranks 11–20. Among the pages it consulted, rank alone explains 29% of the choice; adding every page feature we measured lifts that to 33%. Content still matters at equal rank — pages from the brand's own website (×1.58 odds), pages with more headings (×1.31) and shorter pages (×0.66 per standard deviation of length) are cited more; lists (×0.58), comparison pages (×0.33) and Reddit threads (×0.29) are consulted often but cited rarely — but the recommended practices play no role here either.

Finding 4 — The engines are not Google

ChatGPT splits a question into about four keyword-style queries per answer, keeps 43% of the question's words, adds the current year to 24% of its queries and a brand the user did not name to 27%, and often searches a second time for vendors found in the first round. Gemini writes about three natural-language sub-questions closer to the original. Both rewrite the question anew on every answer: two runs on the same day share 2.4% of their queries and 19% of their cited pages.

The pages they retrieve are not Google's. 97% of the pages ChatGPT retrieves are absent from Google's top 50 for the question as asked, and 99% from Bing's top 30; 0.1–0.6% of the pages the engines cite are in Google's top 10. Gemini, grounded on Google Search, cites 5.6% of Google's top 10 — the websites overlap (about half of Gemini's pages are on websites that appear in Google's top 50), the pages do not. The engines rank pages by how well they match the engines' own rewritten queries, in meaning rather than words: a page's rank correlates with its embedding similarity to those queries (ρ = 0.22) far more than with keyword overlap (ρ = 0.08).

Links to the page do matter — but before, not after, retrieval. Within a website, one standard deviation more referring domains goes with 22.7% more citations, almost all of it from Gemini (+33.7%; ChatGPT +3.7%), and among the pages ChatGPT has already retrieved, links change the odds of citation by nothing (odds ratio 0.97–1.07).

Finding 5 — The model's memory is in the loop, before and after the search

The retrieval is not independent of the model. We asked each engine's model, with no search, which brands and websites it would recommend for every question. Then we watched what it searched for and cited.

  • When a brand is in ChatGPT's prior for a question, ChatGPT writes it into its search queries 60% of the time; when it is not, 7%.
  • Websites in the prior are ranked 1.40× higher in ChatGPT's retrieved lists and have 1.58× the odds of being cited at equal rank.
  • 47% of ChatGPT's citations come from websites its model names without searching, against a 25% baseline when the prior for another question is used.
  • 46% of the brands ChatGPT mentions in the collected answers are not supported by any cited page. A brand in the model's prior is named with 5.8× the odds at equal citation support.

The controlled experiments show the same prior shaping the answer step itself. The same source is cited in 60% of ChatGPT's answers when it is shown first and 43% when it is shown last (Gemini is position-blind, at 57–59%). Revealing the brand names behind sixty packets of identically-specified products changes the top-three rates 2.1× (ChatGPT) and 3.3× (Gemini) as much as simply repeating the prompt; a brand the model holds in low regard for the question loses 19–33 points of top-three rate when its name is attached to a profile it had ranked first anonymously, and a small vendor gains 16. Within a question the model's preferences are stable — one score per brand predicts 82% (ChatGPT) and 77% (Gemini) of its pairwise rankings — and the order in which products are listed changes its ranking more than anything else we varied.

Pages, in other words, are cited when they are searched; brands are named when they are remembered.

Finding 6 — Citations drift

Splitting the 90 days into six two-week windows: within a window, 63–67% of the within-website variation is repeatable. Two weeks later 40% (ChatGPT), 47% (Gemini) and 22% (Perplexity) is still shared; ten weeks later, 14%, 29% and 0%. The engines regenerate their searches for every answer, their search results change over weeks, and pages enter and leave the retrieved set. A citation count from one week says little about the next month.

The niche market: the same mechanism, a different balance

The B2B dataset is the opposite kind of market — dozens of small vendors few people have heard of. The content findings replicate (title match +26.4% per SD, passing all five tests; FAQ block +22.5% across websites and −1.9% within). The retrieval findings replicate (60% of rank-1 pages cited; 99% of retrieved pages absent from Google's top 50).

What changes is the weight of the website and of the model's knowledge. Website authority has no effect (−1.3% per SD, against +11.0% in the consumer market); the models know the tracked vendors for only 2–16% of the questions (against 61–72% for the large consumer providers), so ChatGPT learns vendor names from its first round of results and searches for them in a second. And a cited third-party page that names the vendor raises the odds that the answer names it by ×39.9 in ChatGPT and ×17.7 in Gemini — against ×4.2 and ×6.3 in the consumer market. For a brand the model does not know, being retrievable and being named on third-party pages that are retrievable is not one route to visibility; it is the only one.

What this means for a brand

Four conclusions follow directly from the data, and they are the opposite of most of the advice in circulation.

  1. Write pages for the questions people ask. A title and a paragraph that answer a question in its own terms are the only page properties that reliably earn citations, and they are worth the most on third-party pages.
  2. Being retrieved matters more than anything on the page. The engines search with their own rewritten queries, often naming brands, products and the current year, and rank pages by how well they match those queries in meaning. A good rank in classic search is not a proxy. Replaying the questions through the engines' APIs is the only way to see which pages are retrieved, at what rank, and which are cited.
  3. The recommended AI-optimization practices are not associated with citations within a website. FAQ blocks, schema.org markup, summary boxes, author bylines and page speed may serve readers; our data give no reason to expect citations from them.
  4. A position is a property of the question, the competitors and the engine's retrieval — not of the brand alone. It drifts in weeks and has to be monitored per question, over time.

And two routes to visibility, depending on what the model already knows. For a brand the model knows, visibility lives largely in the model's memory: it is searched for by name, its sites are ranked higher and cited more, and its standing on identical facts is set by what the model already believes — which means the durable work is on the prior, not the page. For a brand the model does not know, there is no stable position without evidence the search can find: pages that match the engines' queries, and third-party pages that name the brand, retrieved together.

This is the empirical basis for how we work at ktau: diagnose the standing per question on the live engines, treat the model's memory and live AI search as two channels that have to be engineered together, and verify the result on the live model over time rather than in a snapshot.

Limitations

The results are associations from a within-website design; they remove everything constant within a website but not differences between pages of the same website that we did not measure, and page links were measured after the citations. Pages were downloaded once, at the end of the collection, and about a third of cited URLs could not be analysed because they were blocked, unavailable or not HTML. The replay uses the engines' APIs, which need not behave like their consumer applications, and Perplexity could not be replayed. The controlled experiments use the small models behind the engines with constructed inputs, and the niche dataset is small. Both datasets are English questions about the US market, written for the organisations that use our monitoring platform rather than sampled from user logs. The engines and their models change continuously; the measurements describe the summer of 2026, and the released logs let later studies test whether the mechanisms persist.

The question sets, citation logs, replay logs and code are released with the paper. Methodology questions, or want to see your category replayed through the engines? Talk to the ktau team.