We followed 158,099 AI answers and 468,045 citations for 90 days across ChatGPT, Gemini and Perplexity, then replayed every question through the engines' own APIs. FAQ blocks, schema markup, summary boxes and author bylines stop mattering the moment you compare pages of the same website. What decides a citation is whether the engine's own search retrieves you — and whether the model already knows your brand.
Generative search engines now decide which web pages a buyer sees as sources. A growing industry tells publishers how to earn those citations: add an FAQ block, add structured data, add a summary box, add an author, make the page faster. Almost none of that advice has been tested outside a controlled setting, and the observational comparisons that circulate are easy to mislead — the practices are concentrated on websites that are cited for reasons that have nothing to do with the practice.
So we tested it properly. Over 90 days we collected every answer the three engines gave to 651 commercial questions in a US consumer market (home internet) and 55 questions in a niche B2B software market, resolved every cited URL to its page and its website, downloaded the pages, and then replayed every question through the engines' own APIs to see the searches they ran, the pages they ranked, and which of those they cited. Finally, we ran controlled experiments on the models behind the engines with the sources held fixed and one factor changed at a time.
Six findings:
The niche B2B market replicates all of it, with one difference that matters: there, the models barely know the vendors, website authority has no effect at all, and a single third-party page that names the brand raises the odds of being named by 40× in ChatGPT and 18× in Gemini.
Two markets, one method. The main dataset is a US consumer market dominated by a few large brands: 651 questions across customer support, equipment, competitive landscape, speed, pricing, performance and reliability, asked daily from 24 June to 21 September 2026 — 143,249 answers and 396,196 citations from ChatGPT (GPT-5.4 nano with web search), Gemini (Gemini 3.1 Flash Lite with Google Search grounding) and Perplexity (Sonar). The second dataset is a niche B2B market of small vendors of proposal, bid and business-development software for architecture, engineering and construction firms: 55 questions, 14,850 answers, 71,849 citations.
Within-website comparison. A page on a popular website is cited more than a page on an obscure one for reasons that have nothing to do with the page. So every effect is estimated by comparing pages of the same website with each other, with a separate intercept for each of 1,117 websites. That single design decision is what makes the recommended practices disappear.
Five tests, not one. A feature counts as real only if it passes five independent checks: it survives false-discovery control over all fourteen features; it is selected in at least 60% of 200 bootstrap Lasso fits; it holds in a double machine-learning estimate with gradient-boosted trees controlling for every other feature; its partial-effect curve is monotone or single-turn; and it keeps its sign in at least four of five variants of the outcome. Only three features pass: title match, best-paragraph match and word overlap with the questions.
Replay. We sent the 651 questions back through the engines' APIs five times over three days with search enabled, and recorded the search queries each engine wrote, the ranked list of pages it retrieved, and which it cited. For ChatGPT, 100% of the URLs cited in the collection were pages the replay retrieved. The replay also lets us observe the model's prior directly: we ask each engine's model, without search, which brands and websites it would recommend for every question, and compare those to what it later searches for and cites.
Controlled experiments. On the models behind the engines (without search), we fixed a set of sources and varied one factor at a time: the position of a source in the list the model reads, whether the brand and publisher names are shown or masked, and whether a product is added or removed from the set — 60 packets of five competing products described by identical fixed profiles, 1,920 answers.
Percent more citations for a page with the practice versus one without. None of the within-website estimates is statistically significant.
The explanation is Simpson's paradox. FAQ blocks appear on 43% of the pages of the most-cited quarter of websites and on 23% of the pages of the least-cited quarter. Compare across websites and you credit the FAQ block with the website's reputation, size and topical focus. Compare within websites and the credit evaporates. The same pattern holds on third-party pages alone, in the second market, and for every engine separately.
This does not mean the practices are useless — an FAQ block may serve readers or classic search perfectly well. It means our data give no reason to expect citations from AI engines because of them.
The three features that survive all five tests measure one thing: how well the page's text matches the questions people ask. One standard deviation better in the paragraph that best matches the questions earns 15.6% more citations (95% CI 11.2–20.2%); in the title, 14.1%; in plain word overlap, 9.1%. On third-party pages — the ones most publishers can actually influence — the effects grow to 25.7%, 23.8% and 12.9%.
Matching also decides which questions a page is cited for. For every page and every query, the query the page is cited for matches it better than one it is not cited for in 88% of comparisons (within-page AUC 0.878; ChatGPT 0.873, Gemini 0.880, Perplexity 0.894). Length, outbound links, headings, lists and tables have no effect at all; there is no page length that works best.
Of the variation in citation counts between pages of the same website, 72% is stable and repeatable — the same pages win in odd and even days, with a reliability of 0.90 over 90 days. Page features explain 2.8% of it for ChatGPT. The engine's own retrieval explains 27%.
The replay shows why. For each question ChatGPT consults about 26 pages and cites 3. It cites 48.7% of the pages it ranks first, 36.8% of those ranked second, 12.3% of ranks 6–10 and 3.3% of ranks 11–20. Among the pages it consulted, rank alone explains 29% of the choice; adding every page feature we measured lifts that to 33%. Content still matters at equal rank — pages from the brand's own website (×1.58 odds), pages with more headings (×1.31) and shorter pages (×0.66 per standard deviation of length) are cited more; lists (×0.58), comparison pages (×0.33) and Reddit threads (×0.29) are consulted often but cited rarely — but the recommended practices play no role here either.
ChatGPT splits a question into about four keyword-style queries per answer, keeps 43% of the question's words, adds the current year to 24% of its queries and a brand the user did not name to 27%, and often searches a second time for vendors found in the first round. Gemini writes about three natural-language sub-questions closer to the original. Both rewrite the question anew on every answer: two runs on the same day share 2.4% of their queries and 19% of their cited pages.
The pages they retrieve are not Google's. 97% of the pages ChatGPT retrieves are absent from Google's top 50 for the question as asked, and 99% from Bing's top 30; 0.1–0.6% of the pages the engines cite are in Google's top 10. Gemini, grounded on Google Search, cites 5.6% of Google's top 10 — the websites overlap (about half of Gemini's pages are on websites that appear in Google's top 50), the pages do not. The engines rank pages by how well they match the engines' own rewritten queries, in meaning rather than words: a page's rank correlates with its embedding similarity to those queries (ρ = 0.22) far more than with keyword overlap (ρ = 0.08).
Links to the page do matter — but before, not after, retrieval. Within a website, one standard deviation more referring domains goes with 22.7% more citations, almost all of it from Gemini (+33.7%; ChatGPT +3.7%), and among the pages ChatGPT has already retrieved, links change the odds of citation by nothing (odds ratio 0.97–1.07).
The retrieval is not independent of the model. We asked each engine's model, with no search, which brands and websites it would recommend for every question. Then we watched what it searched for and cited.
The controlled experiments show the same prior shaping the answer step itself. The same source is cited in 60% of ChatGPT's answers when it is shown first and 43% when it is shown last (Gemini is position-blind, at 57–59%). Revealing the brand names behind sixty packets of identically-specified products changes the top-three rates 2.1× (ChatGPT) and 3.3× (Gemini) as much as simply repeating the prompt; a brand the model holds in low regard for the question loses 19–33 points of top-three rate when its name is attached to a profile it had ranked first anonymously, and a small vendor gains 16. Within a question the model's preferences are stable — one score per brand predicts 82% (ChatGPT) and 77% (Gemini) of its pairwise rankings — and the order in which products are listed changes its ranking more than anything else we varied.
Pages, in other words, are cited when they are searched; brands are named when they are remembered.
Splitting the 90 days into six two-week windows: within a window, 63–67% of the within-website variation is repeatable. Two weeks later 40% (ChatGPT), 47% (Gemini) and 22% (Perplexity) is still shared; ten weeks later, 14%, 29% and 0%. The engines regenerate their searches for every answer, their search results change over weeks, and pages enter and leave the retrieved set. A citation count from one week says little about the next month.
The B2B dataset is the opposite kind of market — dozens of small vendors few people have heard of. The content findings replicate (title match +26.4% per SD, passing all five tests; FAQ block +22.5% across websites and −1.9% within). The retrieval findings replicate (60% of rank-1 pages cited; 99% of retrieved pages absent from Google's top 50).
What changes is the weight of the website and of the model's knowledge. Website authority has no effect (−1.3% per SD, against +11.0% in the consumer market); the models know the tracked vendors for only 2–16% of the questions (against 61–72% for the large consumer providers), so ChatGPT learns vendor names from its first round of results and searches for them in a second. And a cited third-party page that names the vendor raises the odds that the answer names it by ×39.9 in ChatGPT and ×17.7 in Gemini — against ×4.2 and ×6.3 in the consumer market. For a brand the model does not know, being retrievable and being named on third-party pages that are retrievable is not one route to visibility; it is the only one.
Four conclusions follow directly from the data, and they are the opposite of most of the advice in circulation.
And two routes to visibility, depending on what the model already knows. For a brand the model knows, visibility lives largely in the model's memory: it is searched for by name, its sites are ranked higher and cited more, and its standing on identical facts is set by what the model already believes — which means the durable work is on the prior, not the page. For a brand the model does not know, there is no stable position without evidence the search can find: pages that match the engines' queries, and third-party pages that name the brand, retrieved together.
This is the empirical basis for how we work at ktau: diagnose the standing per question on the live engines, treat the model's memory and live AI search as two channels that have to be engineered together, and verify the result on the live model over time rather than in a snapshot.
The results are associations from a within-website design; they remove everything constant within a website but not differences between pages of the same website that we did not measure, and page links were measured after the citations. Pages were downloaded once, at the end of the collection, and about a third of cited URLs could not be analysed because they were blocked, unavailable or not HTML. The replay uses the engines' APIs, which need not behave like their consumer applications, and Perplexity could not be replayed. The controlled experiments use the small models behind the engines with constructed inputs, and the niche dataset is small. Both datasets are English questions about the US market, written for the organisations that use our monitoring platform rather than sampled from user logs. The engines and their models change continuously; the measurements describe the summer of 2026, and the released logs let later studies test whether the mechanisms persist.
The question sets, citation logs, replay logs and code are released with the paper. Methodology questions, or want to see your category replayed through the engines? Talk to the ktau team.