Ask an AI engine a question. Ask again five minutes later. About two thirds of the sources it cites will have changed, and nothing about your site moved at all.
The instrument moves more than the thing it measures
Every chapter so far has deferred one question. How would you know any of this worked?
It is the question a board asks first, and it is the one this whole discipline answers worst.
Here is our answer. You will not like the first half.
In April 2026 three researchers published the measurement study this field needed. Four engines, tracked for forty-six days. ChatGPT, Perplexity, Gemini, Google AI Mode.
They compared the sources cited on one day against the sources cited the next. Then they ran the same comparison inside a single day. That second test is the one that matters.
The day-to-day overlap averaged between 0.34 and 0.42. That is the whole chapter in one figure.
Read that as a percentage and it means roughly 65% of cited sources change overnight. Two thirds of your measurement, gone by morning.
Rank order was worse: 0.21 to 0.26. Not only do the sources change, so does the order they appear in. There goes rank tracking.
Now the finding that decides everything.
"the actual pairwise Jaccard similarity for sources averages between 0.32 and 0.43 across campaigns, values in the same range as the day-to-day figures in Section 4, confirming that intra-day stochastic variation alone accounts for most of the observed instability."
Read that twice. Inside one day, the churn is the same as across days.
So it is not news moving. It is not your competitors shipping. It is not an algorithm update you missed. It is the machine being a machine.
Their own sentence for practitioners is blunter than ours would be. Query an engine once, they say, and the snapshot "may differ substantially from a second query executed minutes later under identical conditions." Minutes.
Under identical conditions.
That is the phrase to keep, and it is theirs, not ours.
One more number from the same study, and it breaks most dashboards quietly. ChatGPT only runs a web search for some questions. In their sample, 57.8% of runs returned zero citations.
Think about what that does to a percentage. If more than half your runs cite nothing, what is the denominator on your citation share?
Most tools will not tell you. Some do not know. Ask anyway.
What one answer is worth
A second paper, July 2026, asked a narrower question. How much of what an AI says about your brand is actually about your brand?
The design was large: 12,933 responses, 20 brands, 8 languages, 3 models.
Then they decomposed the variance. When an answer differs, what actually drove the difference?
The language of the query explained 26.5% of the variation in a single response.
Brand identity explained 1.5%. One and a half percent of the variation.
Sit with those two numbers side by side. What language you asked in mattered seventeen times more than which brand you asked about.
Their conclusion is flat: "a single AI answer carries almost no brand-discriminating signal."
Reliability for brand ranking came out near 0.01 on one answer. Near zero. It reaches about 0.36 only across the full crossed design of languages and models.
So reliability is bought by spreading, not by repeating. Ask in more places, not the same place more often.
One honest caveat before you quote that at somebody. The outcome measured there was sentiment, not citation share. The shape of the finding transfers. The specific number does not.
So how many times do you have to ask?
Somebody has actually computed this, which surprised us.
The April study asks how many runs are enough. Their answer: at least seven runs per prompt per day for brand visibility, and at least eight when you care which sources were cited.
Seven runs. Per prompt. Every single day.
Is that what you are buying?
Now do the arithmetic on a hundred prompts. Seven hundred queries a day. Daily. You can see why most dashboards skip it.
Then the July paper adds the other half. Repeats have sharply diminishing returns, and it puts a number on that: a repeat past the fifth improves reliability by 0.0003. Which is nothing.
So there is a floor and a ceiling. Below roughly seven runs, you are reading noise. Above roughly five, more runs buy almost nothing. A narrow window.
So spread the budget sideways instead: more prompts, more languages, more engines. Not more repeats.
And do not carry these numbers around as universal. The survey that reviewed this literature says so directly: the figure "derives from a small universe of Swiss queries and at most ten repetitions."
Take the method. Leave the constant behind, because it was measured somewhere else.
Why it moves
Four things happen between your question and an answer. Each one is a coin flip.
First, the engine decides whether to search at all. Sometimes it just answers from memory.
Second, if it searches, retrieval returns a set of documents. That set is not fixed.
Third, only some of those documents fit in the context the model gets. Something is always dropped.
Fourth, the model decides which of what it read to cite. Not everything used gets named.
Four gates, and you sit behind all of them. Most dashboards report the last one. Then they call it visibility. It is not.
The critical survey of this literature writes the chain out properly. Your odds of being cited are the odds a search fires, times the odds you are retrieved given it fired, times the odds you are cited given you were retrieved. Three multiplications. Every one below one.
Which has a consequence worth stating. A tool that computes your share of citations only among answers that had citations has silently dropped the first gate. And the first gate was the biggest one.
The first gate was more than half the runs.
Can you fix this by pinning the temperature to zero? No, and somebody checked.
The survey again. A reported temperature of zero "fixes neither the index, nor retrieval, nor the versions of external services."
The model is one source of randomness. It is not the largest one. The index underneath it is.
For contrast, hold it against classical search. One study of 4,706 queries found month-to-month page overlap of 18% for AI Overviews. Organic Google came in at 45%. Better than double.
Organic results churn too. They churn less than half as much, and you have twenty years of practice reading them.
Do the arithmetic once
Take a number your dashboard might show you. Say it reports 8% citation share on ChatGPT.
Now walk it back through the four gates.
Gate one: did a search fire? In the April study, 57.8% of ChatGPT runs cited nothing at all.
So if that 8% counts only the runs with citations, it describes fewer than half the runs. Across all runs it is closer to 3%.
Same measurement. Different denominator. Nearly triple the number, and nothing about you changed.
Gate two: how many runs produced it? If the answer is one per prompt per day, the published floor is seven.
Gate three: which sources moved? Roughly 65% turn over by tomorrow. Including yours.
Gate four: was your brand named without a link? That does not appear anywhere in the count.
So the honest reading of 8% is not a percentage. It is a range, on an unstated denominator, from a sample size nobody published.
None of that makes the tool useless. It makes the decimal point dishonest.
Report it as a band and you are being accurate. Report it as 8.0% and you are being precise about noise.
What changed in 2026
Now the good news, and there is some: two engines started reporting.
In February 2026 Microsoft shipped AI Performance in Bing Webmaster Tools. Its own framing: "For the first time, you can understand how often your content is cited in generative answers, with clear visibility into which URLs are referenced." First time is right.
Actual citations. Actual URLs. First-party, from the engine that served the answer to a real person.
Then in June 2026 Google launched generative AI performance reports in Search Console.
Two engines, four months apart, both conceding that publishers deserve numbers. That is a real change and this chapter will not undersell it.
Now read the fine print, because both reports are narrower than the headlines.
| Engine | What you get | What you do not get | Since |
|---|---|---|---|
| Impressions in AI Overviews and AI Mode, by page, country and date | Clicks. Any query dimension. AI Overviews separated from AI Mode. Access, unless you are in the rollout | 3 June 2026 | |
| Microsoft | Citation counts, the URLs cited, page-level activity, and a sample of grounding queries | Placement or prominence, "without indicating placement or presentation within a specific answer." The sampling method behind the queries | 10 February 2026 |
| OpenAI | A UTM parameter on outbound links | Any impression, citation or appearance data at all | Not offered |
| Anthropic | Nothing | Publisher documentation covers crawling and blocking only | Not offered |
| Perplexity | Nothing | Webmaster documentation covers crawler control and IP verification only | Not offered |
Take Google's report first. It reports impressions. Google defines them plainly: "Impressions are how many times links to your site were shown to a user in a generative AI feature on Google Search."
Useful. Also the only metric in the report. There is no click number, and no position.
You can group the data by page, by country and by date. There is no query dimension.
So you can learn that a page was shown. You cannot learn what anybody asked. The question stays hidden.
Two more limits worth knowing before you go looking for the tab. Google does not split AI Overviews from AI Mode. The report covers both in one bucket. And it is not everywhere yet, being rolled out in Google's words "to a subset of website owners." Check before you look.
Bing goes further on substance and is franker about its gaps.
It reports total citations, the pages cited, and the key phrases the AI used when retrieving your content. That last one is the closest thing to a query anybody publishes.
Then Microsoft says this about it: "The data shown represents a sample of overall citation activity."
A sample. It does not say how large. It does not say how drawn. You cannot scale it.
Which is more honest than most vendors manage. It also leaves you unable to scale the number up. How would you?
Both reports are worth turning on. Neither is a measurement system. They are two windows into two engines, and there are more engines than two. Five, at least.
What an impression is not
Read Google's definition once more, slowly.
Impressions count "how many times links to your site were shown." Links. To your site.
So what happens when a model names you and links nobody?
Nothing happens. It does not appear at all. Not in Google's report. And not in Bing's, which counts citations "displayed as sources."
Both instruments count the same object: a link. And Chapter 8 spent itself on the case that accumulated mentions are what a model learns you from.
Which leaves the most valuable outcome unmeasured by design. A model says your name, recommends you, describes what you do, and cites somebody else's review page as the source.
You were the answer. Somebody else was the citation. Our clients hit this constantly.
Neither first-party report will ever show you that. Nor will most vendor tools, which key on domains.
So keep two questions apart from here on. Was my page cited? Was my brand named?
They are different questions with different answers. Only one has an instrument.
Cited is not the same as correct
One more thing a citation does not tell you.
An early study of four answer engines checked whether cited sources actually supported the sentences attached to them. Only 51.5% were fully supported.
That work is from 2023 and the engines have improved. The shape of the risk has not changed.
Being cited says a system reached for you. It does not say the system read you correctly, and it certainly does not say the sentence next to your name is true.
Which is worth an hour each quarter. Ask the four big engines about your own category, read what comes back, and check whether the claims attached to you are ones you would make.
That check is manual. It is unscalable and irreplaceable. No dashboard does it for you.
The referral you cannot count
Traffic should be the easy part. It is not.
In July 2025 Cloudflare published crawl-to-referral ratios per AI platform. The spread across one week was six orders of magnitude. At one end, about 70,900 crawls for every referral. At the other, ten referrals per crawl. Same week.
Then it added a caveat that undoes more than it admits.
"However, traffic referred by Claude's native app does not include a Referer: header, and we believe that the same holds true for traffic generated from other native apps as well. As such, because the referral counts only include traffic from the Web-based tools from these providers, these calculations may overstate the respective ratios, but it is unclear by how much."
Read the last clause again. It is unclear by how much.
That is the company sitting on a fifth of the web. Saying it cannot size its own blind spot.
So the mechanism is simple and the consequence is not. A person reads about you in a phone app, opens your site later, and arrives with no referrer. Your analytics files them under direct.
They are not direct at all. They came from an answer you never saw, about a question you will never know.
OpenAI mitigates this in one place. Links rendered in the ChatGPT web product carry utm_source=chatgpt.com. A convention, not a commitment.
No engine guarantees a referrer. Not one has published such a commitment.
So treat every AI referral number you own as a floor. Never as a count.
What the vendors say in their own documentation
We expected to have to argue this section. We did not, because two of the largest vendors have already written it down.
Ahrefs publishes a methodology page for Brand Radar, and it is the most candid document in this category. Start with what it discloses. Prompts run "through the free, publicly available web interfaces" of the engines, not through an API. It lists monthly query volumes per platform too.
That distinction matters more than it looks. An API answer is not the product answer. Different system prompt. Different retrieval. Different everything.
Semrush says the same about itself. Their words: "Prompt responses are captured from real requests and not via any APIs of LLMs."
Credit both. It is the single most load-bearing fact about how a vendor collects, and most never mention it. Does yours?
Then Ahrefs says something we have not seen another vendor say.
"Estimated Impressions weight mentions by Google search volume to model potential exposure. This is a modeling choice, not a measured relationship: we don't claim a validated link between Google search volume and how often a query is asked inside an AI tool."
Read what that is. A vendor telling you its flagship number is modelled, and naming the assumption underneath. And it goes further, calling its metrics "directional indicators, not exact traffic counts."
Semrush concedes the general case on its own help page. Their sentence: "AI search and LLM responses are fast-changing and highly personalized, which means no platform can provide exact numbers on visibility."
No platform can. Their words, about their own category, published on their own help page.
So this chapter is not the one making that claim. Two of the biggest vendors got there first. They wrote it in their own documentation, where almost nobody reads.
One more Ahrefs disclosure, because it changes how you read any citation count: it does not filter hallucinated links. Their reasoning is defensible, that malformed output "reflects real model output."
But it means a citation in your dashboard may point at a URL that never existed. Have you checked?
| Vendor | Live product or API? | Prompt set | Runs per prompt |
|---|---|---|---|
| Ahrefs | States: public web interfaces | Sourcing disclosed, set not published | Not disclosed |
| Semrush | States: real requests, not APIs | Clickstream sourced, set not published | Not disclosed |
| Profound | Not stated | Not published | Daily, count not disclosed |
| Peec AI | Not stated | You supply your own | 24-hour cycle, count not disclosed |
| Others we checked | No published methodology page found | Not published | Not published |
Traced to source
Five numbers you will be quoted. Here is where each one actually comes from.
| The claim you hear | Where it actually comes from | Holds up? |
|---|---|---|
| "GEO lifts visibility by up to 40%" | The foundational 2023 paper, measured in a simulator where the source is already inside a five-document context. No clicks, referrals or traffic observed | Rejected as a general claim |
| "Our client got 5.7x from AEO" | The one controlled field study: treated pages grew 5.7x, untreated pages on the same domain grew 3.5x. Most of it was the platform | Mostly tailwind |
| "Track your rank in ChatGPT" | Nowhere. There is no ranked list to hold a position in, which Chapter 1 established and this chapter measures | No such object |
| "AI referral traffic is about 1% of visits" | No primary source we could locate. It circulates between blogs that cite each other | Unsourced |
| "Citation share is up 30% this quarter" | A measurement whose sources turn over roughly 65% overnight, on a denominator that may exclude the runs that cited nobody | Inside the noise |
Row two is the more interesting one, because somebody finally ran the experiment.
A June 2026 field study took one high-traffic domain. It applied a defined bundle of optimizations to a subset of pages. Then it used the untouched remainder of the same site as a control.
That design is right. The untreated pages absorb the tailwind. What is left over is yours.
Treated pages grew 5.7 times. Untreated pages on the same site grew 3.5 times. Read those together.
Their interrupted time-series estimate of the real effect: 1.82 times.
So the intervention did something. The headline number was mostly the tide coming in. And then the authors do the thing almost nobody does: they run a placebo test on their own result, and report that it fails at p equals 0.16.
Their word for their own finding is "suggestive, not conclusive."
That is the best evidence this field currently has that GEO work moves anything. One domain. One bundle. One outcome that is referral traffic rather than citations, and a placebo test the authors could not clear.
The effect is probably real. You should still know that is the state of the evidence before you build a business case on it. All of it.
The four things you can actually count
Enough demolition. Here is what survives contact with the evidence, and it is shorter than you want.
All four are first-party. None needs a subscription. And none tells you what you wish it told you.
- Crawler hits, by verified user agent, from your own server logs. Split GPTBot from OAI-SearchBot from ChatGPT-User, because they do different jobs and only one of them is search. Verify against the published IP ranges rather than trusting the string, since a user agent is a header anybody can type. This tells you that you are being read. It does not tell you that you are being used.
- Referral sessions where a referrer or a UTM survives. Count them, then write the word FLOOR next to the number and never remove it. Native apps send nothing, so this count is structurally low by an amount nobody, including Cloudflare, can size.
- The first-party engine reports, both of them. Google impressions for AI Overviews and AI Mode. Bing citations and cited URLs. Turn both on, accept that they cover two engines, and do not add them together into one number, because they count different things.
- The off-site corpus counts from Chapter 9. Threads you did not start, videos you did not pay for, independent articles. These move slowly and are not gameable in a week, which is exactly what makes them worth tracking.
The one number worth arguing about
If you keep only one thing from this chapter, keep the ratio.
Count the crawler hits you get from AI operators. Then count the referral sessions they send you.
Divide the first by the second. That is your crawl-to-refer ratio.
Then write it down somewhere you will look again.
Cloudflare publishes the same measure across its network, so you have a benchmark. And the spread there was enormous: about 70,900 to one at the extreme.
Why does this one matter more than a visibility score?
Because both halves are yours. Both come from your own logs. Neither depends on a vendor prompt set, a sampling method you cannot see, or an engine choosing to report.
It also asks the commercial question directly. How much are you being read, against how much are you being sent?
A ratio that widens means more extraction and less return. That is a real business fact, and it is measurable today.
It will not tell you whether your content is good. It will tell you whether the trade is worth making.
What this costs
Almost nothing, which is the uncomfortable part.
The four counts above are a half day of setup and an afternoon per quarter. Log parsing you already have. Two engine reports you switch on. Two searches you run by hand.
The visibility subscription is the line item to examine. They run from a few hundred to several thousand a month. The honest question is not whether the tool works.
It is which decision the number changes. Anything?
If the answer is a slide, you are buying a slide. If it is a prioritization you would otherwise get wrong, that may be worth real money. Which is it?
Ask the runs-per-prompt question first. The answer tells you whether you are buying a distribution or a snapshot.
If you are already doing it
Three quick corrections for teams already tracking something. All free.
Stop reporting a single number. Report a range, or report the count of runs behind it. A visibility score with no distribution is a point estimate. On a distribution the researchers measured as unstable within minutes.
Check your denominator. If more than half of ChatGPT runs cite nothing, then a share of cited runs and a share of all runs are wildly different numbers. With the same name on the chart.
Then separate the two claims you are making. That you appeared is one claim. That appearing did anything is a second claim. And no published study has established that one under experimental control.
What to tell your board
This chapter creates a problem you have to manage upward. So here is the script.
Do not say measurement is impossible. It is not, and it sounds like hedging.
Say this instead: the category has two engines reporting first-party data, both of them partial, and a research base that says single-point measurements are unreliable. Then show the four counts.
Give them the direction and the confidence separately. Crawler hits and off-site counts are solid, so state them flatly. Referral numbers are floors. Label them. Engine reports cover two engines, so name which two.
And put a number on what you do not know. Not a guess, a gap: three of the five engines that matter publish nothing to publishers at all.
Then the sentence that will save you a quarter of grief. Appearing and mattering are two claims, and the second one has never been demonstrated under experimental control by anybody, including your competitors.
Somebody in that room is being sold a dashboard that implies otherwise. Better they hear it from you.
Two more, if you have budget and patience.
Log the full distribution, not the daily winner. If your tool exposes individual runs, keep them. A month of runs answers questions a month of averages cannot.
Then run one honest control. Leave a comparable set of pages untouched for a quarter. It is the cheapest experiment in this book and almost nobody runs it, which is why the evidence base looks the way it does.
The objections this chapter has to answer
"You spent a chapter saying measurement is impossible, then sold a measurement stack."
Not impossible. Just narrower than advertised. The four counts are real, first-party and cheap. None of them is a visibility score, which is the thing this chapter says cannot be delivered.
"Our vendor's numbers move consistently with our campaigns."
They might. They would also move without your campaigns, which is the entire point of the June 2026 study. Treated pages grew 5.7 times. Untreated pages on the same domain grew 3.5 times. Without a control you cannot tell those apart, and almost no case study has one.
"The academic work is preprints and small samples."
Some of it, yes. But the volatility study is four engines over forty-six days, the variance study is 12,933 responses, and the survey covers forty-five studies. It is the best evidence available. If you have better, we would like to read it.
"You are describing August 2026. This will change."
It already has, twice this year. Bing shipped citation reporting in February and Google shipped impressions in June. Re-read Table 11.2 before you quote it.
Then the disclosure. We sell GEO services, and this chapter tells you the category's central metric cannot be measured. That is either integrity or a strange sales strategy. Judge it by whether the four counts work when you run them.
What would change our mind
Three things, in order of how much they would matter.
A randomized study. Same site, matched pages, one variable, citations as the outcome. The June 2026 design is close and measures referral traffic instead.
An engine publishing clicks. Google gives impressions. Bing gives citations. Neither closes the loop to a visit, and one of them could.
A vendor publishing its prompt set and its runs per prompt. One doing it forces the rest. The category improves in a quarter.
What is left
So the ledger for measurement, in four lines:
The instrument is less stable than anything it measures. And the instability is stochastic, not seasonal. Not fixable by you.
Two engines now report first-party, one with impressions and one with citations, and three report nothing.
Every referral number you hold is a floor. Native apps send no referrer, and nobody can size the gap.
The causal claim underneath this entire industry rests on one controlled study. Its own placebo test does not clear significance.
Which is the honest position. It is not a comfortable one to end a book on, so we will not end here.
Eleven chapters have handed you a mechanism, a set of artifacts, and a measurement stack you can defend. What is missing is an order to do them in. Just an order.
That is Chapter 12, and it is ninety days long.
- Turn on both engine reports today. Google gives impressions in AI Overviews and AI Mode, Bing gives citations and cited URLs. Free, first-party, covering two engines only.
- Label every referral number a floor. Native AI apps send no referrer, and the company measuring a fifth of the web says it cannot size what that hides.
- Ask your vendor its runs per prompt. The published floor is seven runs daily. Fewer than that and the number you are shown is a coin flip with a decimal point.
- Check the denominator on any share. More than half of ChatGPT runs in one study cited nothing at all, which moves the same percentage enormously.
- Read your crawler logs by verified user agent. Split search bots from training bots, verify against published IP ranges, and treat the string itself as unreliable.
- Never report appearance and effect as one claim. That you were cited is measurable. That being cited did anything has never been shown under experimental control.
- Schulte, Bleeker and Kaufmann, "Don't Measure Once: Measuring Visibility in AI Search (GEO)," arXiv:2604.07585, submitted 8 April 2026. Four engines (ChatGPT, Perplexity, Gemini, Google AI Mode) over a 45 to 46 day window, 24 January to 20 March 2026, with 4,044 consecutive-day pairs and 3,409 pairwise source comparisons, up to ten runs per engine and prompt group, scored with Jaccard similarity and rank-biased overlap. "the day-to-day Jaccard similarity for cited sources averages between 0.34 and 0.42." / "The RBO scores are consistently lower than Jaccard (0.21-0.26), indicating that not only do the source sets change, but so does the rank order in which they appear." / "the actual pairwise Jaccard similarity for sources averages between 0.32 and 0.43 across campaigns, values in the same range as the day-to-day figures in Section 4, confirming that intra-day stochastic variation alone accounts for most of the observed instability." / "if a marketer queries an AI search engine once on a given day, the resulting brand-visibility snapshot may differ substantially from a second query executed minutes later under identical conditions." / "ChatGPT activates web search only for specific queries, leaving 57.8% of its runs with zero citations." / "Practitioners should therefore use at least 7 runs per prompt per day for brand visibility monitoring, and at least 8 runs when source-level coverage matters." / "Brand-level day-to-day stability (Jaccard 0.45-0.59) exceeds source-level stability (0.34-0.42)." Preprint, not peer reviewed at time of writing. https://arxiv.org/abs/2604.07585
- Zatuchin, arXiv:2607.13304, submitted 14 July 2026. A fully crossed corpus of 12,933 responses across 20 Central and Eastern European brands, 8 languages and 3 models, with a stability subset of 1,435 cells resampled about five times. "Query language is the largest systematic facet (26.5% of the variance of one response) against 1.5% for brand identity (ICC 0.0146), so a single AI answer carries almost no brand-discriminating signal." / "Brand-ranking reliability stays low, near 0.01 for a single answer and about 0.36 at the full crossed design, so reliability is bought by spreading across languages and models, not by repeating one prompt." / "a repeat past the fifth reduces it by only 0.0003." Important scope limit: "The outcome is per-response multilingual sentiment polarity," not citation share. Preprint. https://arxiv.org/abs/2607.13304
- Martinez, "Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)," arXiv:2607.14035, 15 July 2026. Reviews 45 studies under a November 2023 to July 2026 window. Its closing judgement: "the evidence is narrow: already-retrieved content can causally alter its citation or use, but no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior." On the foundational GEO paper's headline figure: "The up to 40% figure from the foundational paper is often recast as a general promise of ranking highly in ChatGPT, although it describes a relative visibility gain in a simulator in which five documents have already been placed in context," and the survey formally rates the general claim "Rejected." On reproducibility: "A generative engine is repeatable as an experiment only at the distributional level. Even a reported temperature of zero fixes neither the index, nor retrieval, nor the versions of external services." On the seven-runs figure: "This number is not a universal standard: it derives from a small universe of Swiss queries and at most ten repetitions." It also sets out the decomposition used in this chapter, that the probability of citation is the probability search activates, times the probability of retrieval given activation, times the probability of citation given retrieval, and warns that "A dashboard cannot calculate a share of citations only among responses that contain citations and then interpret it as overall visibility." https://arxiv.org/abs/2607.14035
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, "GEO: Generative Engine Optimization," KDD 2024, pages 5 to 16, arXiv:2311.09735. Source of the widely repeated claim that GEO "can boost visibility by up to 40% in generative engine responses." The measurement setting is a simulator in which the candidate source is already present in a five-document context, and the study observes no clicks, referrals, traffic or purchases. https://arxiv.org/abs/2311.09735
- Watanabe and Nakayashiki, "Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic," arXiv:2606.04362, submitted 3 June 2026. A single high-traffic domain, a defined bundle of interventions in January 2026 applied to one subset of pages, with the untreated remainder of the same domain as a contemporaneous control. "on monthly aggregates total ChatGPT referrals grew 5.7x while untreated pages on the same domain grew 3.5x over the same window." / "an interrupted time-series model on the weekly treated/control ratio estimates a discrete, intervention-aligned level increase of 1.82x (95% CI 1.31-2.54, HAC p=0.001), however, a conservative placebo-in-time permutation test yields p=0.16, so the effect is suggestive, not conclusive, given a short and noisy pre-period." The outcome measured is referral traffic, not citations. Preprint. https://arxiv.org/abs/2606.04362
- Liu, Zhang and Liang, "Evaluating Verifiability in Generative Search Engines," Findings of EMNLP 2023. "on average, a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence." The engines evaluated were Bing Chat, NeevaAI, Perplexity and YouChat, and all have changed substantially since. https://aclanthology.org/2023.findings-emnlp.467/
- Google Search Central, "Introducing generative AI performance reports in Search Console," 3 June 2026. "Today, we're excited to announce the launch of new Search Generative AI performance reports in Search Console." The post also states that the AI data remains inside the overall performance report, and that the new view is an additional dedicated view rather than a removal. https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports
- Google, Search Console Help, generative AI performance report (Search), retrieved 20 August 2026. "The generative AI performance report includes impressions for the following generative AI capabilities on Google Search:" followed by a two-item list, AI Overviews and AI Mode. / "Impressions are how many times links to your site were shown to a user in a generative AI feature on Google Search." / Dimensions offered are Pages, Countries and Dates; the help text lists no queries dimension. / "We're rolling out this report to a subset of website owners, allowing for thorough testing before rolling it further." / "Search Console doesn't include data from experiments in Search Labs, as these experiments are still in active development." https://support.google.com/webmasters/answer/16984139
- Microsoft, "Introducing AI Performance in Bing Webmaster Tools," public preview, 10 February 2026. "For the first time, you can understand how often your content is cited in generative answers, with clear visibility into which URLs are referenced and how citation activity changes over time." / On total citations: "This highlights how often your content is referenced by AI systems, without indicating placement or presentation within a specific answer." / On grounding queries: "Shows the key phrases the AI used when retrieving content that was referenced in AI-generated answers. The data shown represents a sample of overall citation activity." No sampling method is given. / "This release is an early step toward Generative Engine Optimization (GEO) tooling in Bing Webmaster Tools." https://blogs.bing.com/webmaster/february-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview
- OpenAI, publishers and developers FAQ, retrieved 20 August 2026. Documents a UTM convention on outbound links, "ChatGPT automatically includes the UTM parameter utm_source=chatgpt.com in referral URLs." No impression, citation or appearance reporting is offered. Anthropic's publisher-facing documentation covers crawling and blocking only, and Perplexity's covers crawler control and IP verification only. Neither offers publisher analytics of any kind as of 20 August 2026. https://help.openai.com/en/articles/12627856-publishers-and-developers-faq
- Cloudflare, "From Googlebot to GPTBot: the crawl-to-refer ratio on Radar," 1 July 2025. "However, traffic referred by Claude's native app does not include a Referer: header, and we believe that the same holds true for traffic generated from other native apps as well. As such, because the referral counts only include traffic from the Web-based tools from these providers, these calculations may overstate the respective ratios, but it is unclear by how much." / For the week of 19 to 26 June 2025: "the ratios range from Anthropic's 70,900:1 down to Mistral's 0.1:1." / Method: ratios are computed by dividing HTML requests from a platform's crawler user agents by HTML requests carrying that platform's hostname in the Referer header. / "referral traffic coming from Google's ASN (AS15169) is specifically excluded from analysis here" because of prefetching driven by speculation rules. https://blog.cloudflare.com/ai-search-crawl-refer-ratio-on-radar/
- Cloudflare, "Verified bots with cryptography," 1 July 2025. On why user-agent counting alone is unreliable: "Existing identification methods rely on a combination of IP address range (which may be shared by other services, or change over time) and user-agent header (easily spoofable)." https://blog.cloudflare.com/verified-bots-with-cryptography/
- Ahrefs, Brand Radar methodology, updated 26 February 2026. "All prompts run through the free, publicly available web interfaces of ChatGPT, Gemini, Perplexity, Copilot, and other supported platforms to reflect typical user experiences." / "Estimated Impressions weight mentions by Google search volume to model potential exposure. This is a modeling choice, not a measured relationship: we don't claim a validated link between Google search volume and how often a query is asked inside an AI tool." / "Metrics are directional indicators, not exact traffic counts - best understood as modeled visibility signals, and not performance metrics." / "LLMs occasionally generate hallucinated or malformed links. We do not filter out hallucinated or malformed links, as they reflect real model output." Monthly query volumes are disclosed per platform. The prompt set itself is not published. Ahrefs sells the product described. https://ahrefs.com/blog/brand-radar-methodology/
- Semrush, AI visibility data help page, retrieved 20 August 2026. "Prompt responses are captured from real requests and not via any APIs of LLMs." / "AI search and LLM responses are fast-changing and highly personalized, which means no platform can provide exact numbers on visibility." / "We source billions of real prompts from AI search clickstream data and Google's keyword dataset for AI Overviews." Corpus reported as over 317 million prompts and responses across ChatGPT, Gemini, AI Overviews and AI Mode, across 117 regional databases. The prompt set itself is not published. Semrush sells the product described. https://www.semrush.com/kb/1607-semrush-ai-visibility-data
- Profound, answer engine insights documentation, and Peec AI prompt setup documentation, retrieved 20 August 2026. Profound describes sending prompts to answer engines daily without disclosing the set or the runs per prompt, and separately describes an index built on licensed panel conversations. Peec AI runs customer-supplied prompts on a 24-hour cycle without disclosing runs per prompt. For Scrunch, Athena, Conductor and Similarweb's AI products, no public methodology page stating prompt set, sample size, or whether queries run against the live product or an API could be located on 20 August 2026. https://help.tryprofound.com/articles/3443229936-answer-engine-insights-overview and https://docs.peec.ai/setting-up-your-prompts
- Kirsten et al., as reported in the Martinez survey above: 4,706 queries across several surfaces in the United States and Germany, finding two-month page overlap of 18% for AI Overviews against 45% for organic Google, and 9 to 28% of decisions changing on repeated runs where temperature could be set to zero. Findings of ACL 2026. This chapter carries the figure at one remove and has not read the primary. https://aclanthology.org/2026.findings-acl.526/
- The "AI referral traffic is about 1% of visits" figure that circulates widely could not be traced to any primary vendor publication on 20 August 2026. It is therefore not used in this chapter as a number, only as an example of an unsourced claim.