Library/ GEO Playbook 2027/ Chapter 11
Chapter 11 of 12

Measuring AI Visibility: What Can Actually Be Counted

Ask an AI engine a question. Ask again five minutes later. About two thirds of the sources it cites will have changed, and nothing about your site moved at all. The instrument moves more than the thing it measures Every chapter so far has deferred one question. How would you know any of this worked? […]

6 min read Updated Aug 2026 Part IV · The Operation

Ask an AI engine a question. Ask again five minutes later. About two thirds of the sources it cites will have changed, and nothing about your site moved at all.

The instrument moves more than the thing it measures

Every chapter so far has deferred one question. How would you know any of this worked?

It is the question a board asks first, and it is the one this whole discipline answers worst.

Here is our answer. You will not like the first half.

In April 2026 three researchers published the measurement study this field needed. Four engines, tracked for forty-six days. ChatGPT, Perplexity, Gemini, Google AI Mode.

They compared the sources cited on one day against the sources cited the next. Then they ran the same comparison inside a single day. That second test is the one that matters.

The day-to-day overlap averaged between 0.34 and 0.42. That is the whole chapter in one figure.

Read that as a percentage and it means roughly 65% of cited sources change overnight. Two thirds of your measurement, gone by morning.

Rank order was worse: 0.21 to 0.26. Not only do the sources change, so does the order they appear in. There goes rank tracking.

Now the finding that decides everything.

Schulte, Bleeker and Kaufmann, April 2026

"the actual pairwise Jaccard similarity for sources averages between 0.32 and 0.43 across campaigns, values in the same range as the day-to-day figures in Section 4, confirming that intra-day stochastic variation alone accounts for most of the observed instability."

Read that twice. Inside one day, the churn is the same as across days.

So it is not news moving. It is not your competitors shipping. It is not an algorithm update you missed. It is the machine being a machine.

Their own sentence for practitioners is blunter than ours would be. Query an engine once, they say, and the snapshot "may differ substantially from a second query executed minutes later under identical conditions." Minutes.

Under identical conditions.

That is the phrase to keep, and it is theirs, not ours.

One more number from the same study, and it breaks most dashboards quietly. ChatGPT only runs a web search for some questions. In their sample, 57.8% of runs returned zero citations.

Think about what that does to a percentage. If more than half your runs cite nothing, what is the denominator on your citation share?

Most tools will not tell you. Some do not know. Ask anyway.

What one answer is worth

A second paper, July 2026, asked a narrower question. How much of what an AI says about your brand is actually about your brand?

The design was large: 12,933 responses, 20 brands, 8 languages, 3 models.

Then they decomposed the variance. When an answer differs, what actually drove the difference?

The language of the query explained 26.5% of the variation in a single response.

Brand identity explained 1.5%. One and a half percent of the variation.

Sit with those two numbers side by side. What language you asked in mattered seventeen times more than which brand you asked about.

Their conclusion is flat: "a single AI answer carries almost no brand-discriminating signal."

Reliability for brand ranking came out near 0.01 on one answer. Near zero. It reaches about 0.36 only across the full crossed design of languages and models.

So reliability is bought by spreading, not by repeating. Ask in more places, not the same place more often.

One honest caveat before you quote that at somebody. The outcome measured there was sentiment, not citation share. The shape of the finding transfers. The specific number does not.

So how many times do you have to ask?

Somebody has actually computed this, which surprised us.

The April study asks how many runs are enough. Their answer: at least seven runs per prompt per day for brand visibility, and at least eight when you care which sources were cited.

Seven runs. Per prompt. Every single day.

Is that what you are buying?

Now do the arithmetic on a hundred prompts. Seven hundred queries a day. Daily. You can see why most dashboards skip it.

Then the July paper adds the other half. Repeats have sharply diminishing returns, and it puts a number on that: a repeat past the fifth improves reliability by 0.0003. Which is nothing.

So there is a floor and a ceiling. Below roughly seven runs, you are reading noise. Above roughly five, more runs buy almost nothing. A narrow window.

So spread the budget sideways instead: more prompts, more languages, more engines. Not more repeats.

And do not carry these numbers around as universal. The survey that reviewed this literature says so directly: the figure "derives from a small universe of Swiss queries and at most ten repetitions."

Take the method. Leave the constant behind, because it was measured somewhere else.

Figure 11.1 · The churn is not the calendarSource overlap between two measurements of the same prompt
SHARE OF CITED SOURCES THAT ARE THE SAME IN BOTH MEASUREMENTS 0.0 0.5 1.0 One day apart 0.34 to 0.42 Minutes apart 0.32 to 0.43 Rank order, one day 0.21 to 0.26 The top two bars are the finding. Same range, so time is not the cause.
If the day-apart bar were long and the minutes-apart bar were short, the churn would be news cycles and index updates, and you could measure around it by sampling less often. They are the same length. The variation is inside the machine, which means sampling less often does not help and sampling more often is the only thing that does. The third bar is the one that kills rank tracking outright: the order sources appear in is even less stable than which sources appear at all.

Why it moves

Four things happen between your question and an answer. Each one is a coin flip.

First, the engine decides whether to search at all. Sometimes it just answers from memory.

Second, if it searches, retrieval returns a set of documents. That set is not fixed.

Third, only some of those documents fit in the context the model gets. Something is always dropped.

Fourth, the model decides which of what it read to cite. Not everything used gets named.

Four gates, and you sit behind all of them. Most dashboards report the last one. Then they call it visibility. It is not.

The critical survey of this literature writes the chain out properly. Your odds of being cited are the odds a search fires, times the odds you are retrieved given it fired, times the odds you are cited given you were retrieved. Three multiplications. Every one below one.

Which has a consequence worth stating. A tool that computes your share of citations only among answers that had citations has silently dropped the first gate. And the first gate was the biggest one.

The first gate was more than half the runs.

Can you fix this by pinning the temperature to zero? No, and somebody checked.

The survey again. A reported temperature of zero "fixes neither the index, nor retrieval, nor the versions of external services."

The model is one source of randomness. It is not the largest one. The index underneath it is.

For contrast, hold it against classical search. One study of 4,706 queries found month-to-month page overlap of 18% for AI Overviews. Organic Google came in at 45%. Better than double.

Organic results churn too. They churn less than half as much, and you have twenty years of practice reading them.

Do the arithmetic once

Take a number your dashboard might show you. Say it reports 8% citation share on ChatGPT.

Now walk it back through the four gates.

Gate one: did a search fire? In the April study, 57.8% of ChatGPT runs cited nothing at all.

So if that 8% counts only the runs with citations, it describes fewer than half the runs. Across all runs it is closer to 3%.

Same measurement. Different denominator. Nearly triple the number, and nothing about you changed.

Gate two: how many runs produced it? If the answer is one per prompt per day, the published floor is seven.

Gate three: which sources moved? Roughly 65% turn over by tomorrow. Including yours.

Gate four: was your brand named without a link? That does not appear anywhere in the count.

So the honest reading of 8% is not a percentage. It is a range, on an unstated denominator, from a sample size nobody published.

None of that makes the tool useless. It makes the decimal point dishonest.

Report it as a band and you are being accurate. Report it as 8.0% and you are being precise about noise.

What changed in 2026

Now the good news, and there is some: two engines started reporting.

In February 2026 Microsoft shipped AI Performance in Bing Webmaster Tools. Its own framing: "For the first time, you can understand how often your content is cited in generative answers, with clear visibility into which URLs are referenced." First time is right.

Actual citations. Actual URLs. First-party, from the engine that served the answer to a real person.

Then in June 2026 Google launched generative AI performance reports in Search Console.

Two engines, four months apart, both conceding that publishers deserve numbers. That is a real change and this chapter will not undersell it.

Now read the fine print, because both reports are narrower than the headlines.

Table 11.2 · What each engine will actually tell youFirst-party reporting, checked 20 August 2026
EngineWhat you getWhat you do not getSince
GoogleImpressions in AI Overviews and AI Mode, by page, country and dateClicks. Any query dimension. AI Overviews separated from AI Mode. Access, unless you are in the rollout3 June 2026
MicrosoftCitation counts, the URLs cited, page-level activity, and a sample of grounding queriesPlacement or prominence, "without indicating placement or presentation within a specific answer." The sampling method behind the queries10 February 2026
OpenAIA UTM parameter on outbound linksAny impression, citation or appearance data at allNot offered
AnthropicNothingPublisher documentation covers crawling and blocking onlyNot offered
PerplexityNothingWebmaster documentation covers crawler control and IP verification onlyNot offered
Read the two right-hand columns together and the shape is clear. Google gives you impressions and no clicks. Microsoft gives you citations and no placement. Neither gives you the query that produced the answer, so neither closes the loop between what somebody asked and what you got. Note also which way round this is. Bing, the smaller engine, discloses more than Google does, and it did so four months earlier. The three pure AI companies disclose nothing, and two of them have no publisher-facing product to disclose it through.

Take Google's report first. It reports impressions. Google defines them plainly: "Impressions are how many times links to your site were shown to a user in a generative AI feature on Google Search."

Useful. Also the only metric in the report. There is no click number, and no position.

You can group the data by page, by country and by date. There is no query dimension.

So you can learn that a page was shown. You cannot learn what anybody asked. The question stays hidden.

Two more limits worth knowing before you go looking for the tab. Google does not split AI Overviews from AI Mode. The report covers both in one bucket. And it is not everywhere yet, being rolled out in Google's words "to a subset of website owners." Check before you look.

Bing goes further on substance and is franker about its gaps.

It reports total citations, the pages cited, and the key phrases the AI used when retrieving your content. That last one is the closest thing to a query anybody publishes.

Then Microsoft says this about it: "The data shown represents a sample of overall citation activity."

A sample. It does not say how large. It does not say how drawn. You cannot scale it.

Which is more honest than most vendors manage. It also leaves you unable to scale the number up. How would you?

Both reports are worth turning on. Neither is a measurement system. They are two windows into two engines, and there are more engines than two. Five, at least.

What an impression is not

Read Google's definition once more, slowly.

Impressions count "how many times links to your site were shown." Links. To your site.

So what happens when a model names you and links nobody?

Nothing happens. It does not appear at all. Not in Google's report. And not in Bing's, which counts citations "displayed as sources."

Both instruments count the same object: a link. And Chapter 8 spent itself on the case that accumulated mentions are what a model learns you from.

Which leaves the most valuable outcome unmeasured by design. A model says your name, recommends you, describes what you do, and cites somebody else's review page as the source.

You were the answer. Somebody else was the citation. Our clients hit this constantly.

Neither first-party report will ever show you that. Nor will most vendor tools, which key on domains.

So keep two questions apart from here on. Was my page cited? Was my brand named?

They are different questions with different answers. Only one has an instrument.

Cited is not the same as correct

One more thing a citation does not tell you.

An early study of four answer engines checked whether cited sources actually supported the sentences attached to them. Only 51.5% were fully supported.

That work is from 2023 and the engines have improved. The shape of the risk has not changed.

Being cited says a system reached for you. It does not say the system read you correctly, and it certainly does not say the sentence next to your name is true.

Which is worth an hour each quarter. Ask the four big engines about your own category, read what comes back, and check whether the claims attached to you are ones you would make.

That check is manual. It is unscalable and irreplaceable. No dashboard does it for you.

The referral you cannot count

Traffic should be the easy part. It is not.

In July 2025 Cloudflare published crawl-to-referral ratios per AI platform. The spread across one week was six orders of magnitude. At one end, about 70,900 crawls for every referral. At the other, ten referrals per crawl. Same week.

Then it added a caveat that undoes more than it admits.

Cloudflare, 1 July 2025

"However, traffic referred by Claude's native app does not include a Referer: header, and we believe that the same holds true for traffic generated from other native apps as well. As such, because the referral counts only include traffic from the Web-based tools from these providers, these calculations may overstate the respective ratios, but it is unclear by how much."

Read the last clause again. It is unclear by how much.

That is the company sitting on a fifth of the web. Saying it cannot size its own blind spot.

So the mechanism is simple and the consequence is not. A person reads about you in a phone app, opens your site later, and arrives with no referrer. Your analytics files them under direct.

They are not direct at all. They came from an answer you never saw, about a question you will never know.

OpenAI mitigates this in one place. Links rendered in the ChatGPT web product carry utm_source=chatgpt.com. A convention, not a commitment.

No engine guarantees a referrer. Not one has published such a commitment.

So treat every AI referral number you own as a floor. Never as a count.

What the vendors say in their own documentation

We expected to have to argue this section. We did not, because two of the largest vendors have already written it down.

Ahrefs publishes a methodology page for Brand Radar, and it is the most candid document in this category. Start with what it discloses. Prompts run "through the free, publicly available web interfaces" of the engines, not through an API. It lists monthly query volumes per platform too.

That distinction matters more than it looks. An API answer is not the product answer. Different system prompt. Different retrieval. Different everything.

Semrush says the same about itself. Their words: "Prompt responses are captured from real requests and not via any APIs of LLMs."

Credit both. It is the single most load-bearing fact about how a vendor collects, and most never mention it. Does yours?

Then Ahrefs says something we have not seen another vendor say.

Ahrefs, Brand Radar methodology, updated 26 February 2026

"Estimated Impressions weight mentions by Google search volume to model potential exposure. This is a modeling choice, not a measured relationship: we don't claim a validated link between Google search volume and how often a query is asked inside an AI tool."

Read what that is. A vendor telling you its flagship number is modelled, and naming the assumption underneath. And it goes further, calling its metrics "directional indicators, not exact traffic counts."

Semrush concedes the general case on its own help page. Their sentence: "AI search and LLM responses are fast-changing and highly personalized, which means no platform can provide exact numbers on visibility."

No platform can. Their words, about their own category, published on their own help page.

So this chapter is not the one making that claim. Two of the biggest vendors got there first. They wrote it in their own documentation, where almost nobody reads.

One more Ahrefs disclosure, because it changes how you read any citation count: it does not filter hallucinated links. Their reasoning is defensible, that malformed output "reflects real model output."

But it means a citation in your dashboard may point at a URL that never existed. Have you checked?

Table 11.3 · What your vendor disclosesPublished methodology pages, checked 20 August 2026
VendorLive product or API?Prompt setRuns per prompt
AhrefsStates: public web interfacesSourcing disclosed, set not publishedNot disclosed
SemrushStates: real requests, not APIsClickstream sourced, set not publishedNot disclosed
ProfoundNot statedNot publishedDaily, count not disclosed
Peec AINot statedYou supply your own24-hour cycle, count not disclosed
Others we checkedNo published methodology page foundNot publishedNot published
Column four is the one to ask about on your next call, because the April 2026 convergence work puts the floor at seven runs per prompt per day and no vendor in this table publishes its number. That is not an accusation of malpractice. It is a question with a right answer that you are entitled to hear before you renew. Ask it in exactly those words: how many times per day do you run each prompt, and do you report the distribution or the last result?

Traced to source

Five numbers you will be quoted. Here is where each one actually comes from.

Table 11.4 · Five claims, tracedWhat gets repeated, and what the paper says
The claim you hearWhere it actually comes fromHolds up?
"GEO lifts visibility by up to 40%"The foundational 2023 paper, measured in a simulator where the source is already inside a five-document context. No clicks, referrals or traffic observedRejected as a general claim
"Our client got 5.7x from AEO"The one controlled field study: treated pages grew 5.7x, untreated pages on the same domain grew 3.5x. Most of it was the platformMostly tailwind
"Track your rank in ChatGPT"Nowhere. There is no ranked list to hold a position in, which Chapter 1 established and this chapter measuresNo such object
"AI referral traffic is about 1% of visits"No primary source we could locate. It circulates between blogs that cite each otherUnsourced
"Citation share is up 30% this quarter"A measurement whose sources turn over roughly 65% overnight, on a denominator that may exclude the runs that cited nobodyInside the noise
Row one deserves its own sentence because it is the most repeated number in this industry. The paper is real, careful and honest about its setting. It ran in a simulator, with the candidate source already placed in the context window, and it observed no clicks, no traffic and no purchases. A critical survey of forty-five studies rejects the 40% as a general claim for exactly that reason. The finding is not fake. It has been carried a very long way from the room it was measured in.

Row two is the more interesting one, because somebody finally ran the experiment.

A June 2026 field study took one high-traffic domain. It applied a defined bundle of optimizations to a subset of pages. Then it used the untouched remainder of the same site as a control.

That design is right. The untreated pages absorb the tailwind. What is left over is yours.

Treated pages grew 5.7 times. Untreated pages on the same site grew 3.5 times. Read those together.

Their interrupted time-series estimate of the real effect: 1.82 times.

So the intervention did something. The headline number was mostly the tide coming in. And then the authors do the thing almost nobody does: they run a placebo test on their own result, and report that it fails at p equals 0.16.

Their word for their own finding is "suggestive, not conclusive."

That is the best evidence this field currently has that GEO work moves anything. One domain. One bundle. One outcome that is referral traffic rather than citations, and a placebo test the authors could not clear.

The effect is probably real. You should still know that is the state of the evidence before you build a business case on it. All of it.

The four things you can actually count

Enough demolition. Here is what survives contact with the evidence, and it is shorter than you want.

All four are first-party. None needs a subscription. And none tells you what you wish it told you.

Artifact 11.5 · The honest measurement stackFour counts, one page, re-run quarterly
  1. Crawler hits, by verified user agent, from your own server logs. Split GPTBot from OAI-SearchBot from ChatGPT-User, because they do different jobs and only one of them is search. Verify against the published IP ranges rather than trusting the string, since a user agent is a header anybody can type. This tells you that you are being read. It does not tell you that you are being used.
  2. Referral sessions where a referrer or a UTM survives. Count them, then write the word FLOOR next to the number and never remove it. Native apps send nothing, so this count is structurally low by an amount nobody, including Cloudflare, can size.
  3. The first-party engine reports, both of them. Google impressions for AI Overviews and AI Mode. Bing citations and cited URLs. Turn both on, accept that they cover two engines, and do not add them together into one number, because they count different things.
  4. The off-site corpus counts from Chapter 9. Threads you did not start, videos you did not pay for, independent articles. These move slowly and are not gameable in a week, which is exactly what makes them worth tracking.
What it returns. Four numbers and a direction. Re-run them ninety days apart and you have a trend for the two counts that are stable and a floor for the two that are not. That is a smaller claim than a visibility score, and it is one you can defend in a board meeting when somebody asks where the number came from.
What it does not return. Your share of AI answers. Nobody can give you that, and the two vendors most able to try both say so in their own documentation. If you want a prompt-level view anyway, the published floor is seven runs per prompt per day, and any tool running fewer is reporting a coin flip with a decimal point.
Step 1 is the one most teams already have and never look at, because crawler traffic sits in a log nobody reads. It is also the only signal in this list that responds within days of a change you make, which makes it the closest thing to a feedback loop this discipline currently has. Step 4 is the slowest and the most durable. Between them they bracket the question: are the machines reading me, and is there anything about me out there to read.

The one number worth arguing about

If you keep only one thing from this chapter, keep the ratio.

Count the crawler hits you get from AI operators. Then count the referral sessions they send you.

Divide the first by the second. That is your crawl-to-refer ratio.

Then write it down somewhere you will look again.

Cloudflare publishes the same measure across its network, so you have a benchmark. And the spread there was enormous: about 70,900 to one at the extreme.

Why does this one matter more than a visibility score?

Because both halves are yours. Both come from your own logs. Neither depends on a vendor prompt set, a sampling method you cannot see, or an engine choosing to report.

It also asks the commercial question directly. How much are you being read, against how much are you being sent?

A ratio that widens means more extraction and less return. That is a real business fact, and it is measurable today.

It will not tell you whether your content is good. It will tell you whether the trade is worth making.

What this costs

Almost nothing, which is the uncomfortable part.

The four counts above are a half day of setup and an afternoon per quarter. Log parsing you already have. Two engine reports you switch on. Two searches you run by hand.

The visibility subscription is the line item to examine. They run from a few hundred to several thousand a month. The honest question is not whether the tool works.

It is which decision the number changes. Anything?

If the answer is a slide, you are buying a slide. If it is a prioritization you would otherwise get wrong, that may be worth real money. Which is it?

Ask the runs-per-prompt question first. The answer tells you whether you are buying a distribution or a snapshot.

If you are already doing it

Three quick corrections for teams already tracking something. All free.

Stop reporting a single number. Report a range, or report the count of runs behind it. A visibility score with no distribution is a point estimate. On a distribution the researchers measured as unstable within minutes.

Check your denominator. If more than half of ChatGPT runs cite nothing, then a share of cited runs and a share of all runs are wildly different numbers. With the same name on the chart.

Then separate the two claims you are making. That you appeared is one claim. That appearing did anything is a second claim. And no published study has established that one under experimental control.

What to tell your board

This chapter creates a problem you have to manage upward. So here is the script.

Do not say measurement is impossible. It is not, and it sounds like hedging.

Say this instead: the category has two engines reporting first-party data, both of them partial, and a research base that says single-point measurements are unreliable. Then show the four counts.

Give them the direction and the confidence separately. Crawler hits and off-site counts are solid, so state them flatly. Referral numbers are floors. Label them. Engine reports cover two engines, so name which two.

And put a number on what you do not know. Not a guess, a gap: three of the five engines that matter publish nothing to publishers at all.

Then the sentence that will save you a quarter of grief. Appearing and mattering are two claims, and the second one has never been demonstrated under experimental control by anybody, including your competitors.

Somebody in that room is being sold a dashboard that implies otherwise. Better they hear it from you.

Two more, if you have budget and patience.

Log the full distribution, not the daily winner. If your tool exposes individual runs, keep them. A month of runs answers questions a month of averages cannot.

Then run one honest control. Leave a comparable set of pages untouched for a quarter. It is the cheapest experiment in this book and almost nobody runs it, which is why the evidence base looks the way it does.

The objections this chapter has to answer

"You spent a chapter saying measurement is impossible, then sold a measurement stack."

Not impossible. Just narrower than advertised. The four counts are real, first-party and cheap. None of them is a visibility score, which is the thing this chapter says cannot be delivered.

"Our vendor's numbers move consistently with our campaigns."

They might. They would also move without your campaigns, which is the entire point of the June 2026 study. Treated pages grew 5.7 times. Untreated pages on the same domain grew 3.5 times. Without a control you cannot tell those apart, and almost no case study has one.

"The academic work is preprints and small samples."

Some of it, yes. But the volatility study is four engines over forty-six days, the variance study is 12,933 responses, and the survey covers forty-five studies. It is the best evidence available. If you have better, we would like to read it.

"You are describing August 2026. This will change."

It already has, twice this year. Bing shipped citation reporting in February and Google shipped impressions in June. Re-read Table 11.2 before you quote it.

Then the disclosure. We sell GEO services, and this chapter tells you the category's central metric cannot be measured. That is either integrity or a strange sales strategy. Judge it by whether the four counts work when you run them.

What would change our mind

Three things, in order of how much they would matter.

A randomized study. Same site, matched pages, one variable, citations as the outcome. The June 2026 design is close and measures referral traffic instead.

An engine publishing clicks. Google gives impressions. Bing gives citations. Neither closes the loop to a visit, and one of them could.

A vendor publishing its prompt set and its runs per prompt. One doing it forces the rest. The category improves in a quarter.

What is left

So the ledger for measurement, in four lines:

The instrument is less stable than anything it measures. And the instability is stochastic, not seasonal. Not fixable by you.

Two engines now report first-party, one with impressions and one with citations, and three report nothing.

Every referral number you hold is a floor. Native apps send no referrer, and nobody can size the gap.

The causal claim underneath this entire industry rests on one controlled study. Its own placebo test does not clear significance.

Which is the honest position. It is not a comfortable one to end a book on, so we will not end here.

Eleven chapters have handed you a mechanism, a set of artifacts, and a measurement stack you can defend. What is missing is an order to do them in. Just an order.

That is Chapter 12, and it is ninety days long.

Chapter 11 · What to do with this:
  • Turn on both engine reports today. Google gives impressions in AI Overviews and AI Mode, Bing gives citations and cited URLs. Free, first-party, covering two engines only.
  • Label every referral number a floor. Native AI apps send no referrer, and the company measuring a fifth of the web says it cannot size what that hides.
  • Ask your vendor its runs per prompt. The published floor is seven runs daily. Fewer than that and the number you are shown is a coin flip with a decimal point.
  • Check the denominator on any share. More than half of ChatGPT runs in one study cited nothing at all, which moves the same percentage enormously.
  • Read your crawler logs by verified user agent. Split search bots from training bots, verify against published IP ranges, and treat the string itself as unreliable.
  • Never report appearance and effect as one claim. That you were cited is measurable. That being cited did anything has never been shown under experimental control.
Sources
  • Schulte, Bleeker and Kaufmann, "Don't Measure Once: Measuring Visibility in AI Search (GEO)," arXiv:2604.07585, submitted 8 April 2026. Four engines (ChatGPT, Perplexity, Gemini, Google AI Mode) over a 45 to 46 day window, 24 January to 20 March 2026, with 4,044 consecutive-day pairs and 3,409 pairwise source comparisons, up to ten runs per engine and prompt group, scored with Jaccard similarity and rank-biased overlap. "the day-to-day Jaccard similarity for cited sources averages between 0.34 and 0.42." / "The RBO scores are consistently lower than Jaccard (0.21-0.26), indicating that not only do the source sets change, but so does the rank order in which they appear." / "the actual pairwise Jaccard similarity for sources averages between 0.32 and 0.43 across campaigns, values in the same range as the day-to-day figures in Section 4, confirming that intra-day stochastic variation alone accounts for most of the observed instability." / "if a marketer queries an AI search engine once on a given day, the resulting brand-visibility snapshot may differ substantially from a second query executed minutes later under identical conditions." / "ChatGPT activates web search only for specific queries, leaving 57.8% of its runs with zero citations." / "Practitioners should therefore use at least 7 runs per prompt per day for brand visibility monitoring, and at least 8 runs when source-level coverage matters." / "Brand-level day-to-day stability (Jaccard 0.45-0.59) exceeds source-level stability (0.34-0.42)." Preprint, not peer reviewed at time of writing. https://arxiv.org/abs/2604.07585
  • Zatuchin, arXiv:2607.13304, submitted 14 July 2026. A fully crossed corpus of 12,933 responses across 20 Central and Eastern European brands, 8 languages and 3 models, with a stability subset of 1,435 cells resampled about five times. "Query language is the largest systematic facet (26.5% of the variance of one response) against 1.5% for brand identity (ICC 0.0146), so a single AI answer carries almost no brand-discriminating signal." / "Brand-ranking reliability stays low, near 0.01 for a single answer and about 0.36 at the full crossed design, so reliability is bought by spreading across languages and models, not by repeating one prompt." / "a repeat past the fifth reduces it by only 0.0003." Important scope limit: "The outcome is per-response multilingual sentiment polarity," not citation share. Preprint. https://arxiv.org/abs/2607.13304
  • Martinez, "Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)," arXiv:2607.14035, 15 July 2026. Reviews 45 studies under a November 2023 to July 2026 window. Its closing judgement: "the evidence is narrow: already-retrieved content can causally alter its citation or use, but no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior." On the foundational GEO paper's headline figure: "The up to 40% figure from the foundational paper is often recast as a general promise of ranking highly in ChatGPT, although it describes a relative visibility gain in a simulator in which five documents have already been placed in context," and the survey formally rates the general claim "Rejected." On reproducibility: "A generative engine is repeatable as an experiment only at the distributional level. Even a reported temperature of zero fixes neither the index, nor retrieval, nor the versions of external services." On the seven-runs figure: "This number is not a universal standard: it derives from a small universe of Swiss queries and at most ten repetitions." It also sets out the decomposition used in this chapter, that the probability of citation is the probability search activates, times the probability of retrieval given activation, times the probability of citation given retrieval, and warns that "A dashboard cannot calculate a share of citations only among responses that contain citations and then interpret it as overall visibility." https://arxiv.org/abs/2607.14035
  • Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, "GEO: Generative Engine Optimization," KDD 2024, pages 5 to 16, arXiv:2311.09735. Source of the widely repeated claim that GEO "can boost visibility by up to 40% in generative engine responses." The measurement setting is a simulator in which the candidate source is already present in a five-document context, and the study observes no clicks, referrals, traffic or purchases. https://arxiv.org/abs/2311.09735
  • Watanabe and Nakayashiki, "Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic," arXiv:2606.04362, submitted 3 June 2026. A single high-traffic domain, a defined bundle of interventions in January 2026 applied to one subset of pages, with the untreated remainder of the same domain as a contemporaneous control. "on monthly aggregates total ChatGPT referrals grew 5.7x while untreated pages on the same domain grew 3.5x over the same window." / "an interrupted time-series model on the weekly treated/control ratio estimates a discrete, intervention-aligned level increase of 1.82x (95% CI 1.31-2.54, HAC p=0.001), however, a conservative placebo-in-time permutation test yields p=0.16, so the effect is suggestive, not conclusive, given a short and noisy pre-period." The outcome measured is referral traffic, not citations. Preprint. https://arxiv.org/abs/2606.04362
  • Liu, Zhang and Liang, "Evaluating Verifiability in Generative Search Engines," Findings of EMNLP 2023. "on average, a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence." The engines evaluated were Bing Chat, NeevaAI, Perplexity and YouChat, and all have changed substantially since. https://aclanthology.org/2023.findings-emnlp.467/
  • Google Search Central, "Introducing generative AI performance reports in Search Console," 3 June 2026. "Today, we're excited to announce the launch of new Search Generative AI performance reports in Search Console." The post also states that the AI data remains inside the overall performance report, and that the new view is an additional dedicated view rather than a removal. https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports
  • Google, Search Console Help, generative AI performance report (Search), retrieved 20 August 2026. "The generative AI performance report includes impressions for the following generative AI capabilities on Google Search:" followed by a two-item list, AI Overviews and AI Mode. / "Impressions are how many times links to your site were shown to a user in a generative AI feature on Google Search." / Dimensions offered are Pages, Countries and Dates; the help text lists no queries dimension. / "We're rolling out this report to a subset of website owners, allowing for thorough testing before rolling it further." / "Search Console doesn't include data from experiments in Search Labs, as these experiments are still in active development." https://support.google.com/webmasters/answer/16984139
  • Microsoft, "Introducing AI Performance in Bing Webmaster Tools," public preview, 10 February 2026. "For the first time, you can understand how often your content is cited in generative answers, with clear visibility into which URLs are referenced and how citation activity changes over time." / On total citations: "This highlights how often your content is referenced by AI systems, without indicating placement or presentation within a specific answer." / On grounding queries: "Shows the key phrases the AI used when retrieving content that was referenced in AI-generated answers. The data shown represents a sample of overall citation activity." No sampling method is given. / "This release is an early step toward Generative Engine Optimization (GEO) tooling in Bing Webmaster Tools." https://blogs.bing.com/webmaster/february-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview
  • OpenAI, publishers and developers FAQ, retrieved 20 August 2026. Documents a UTM convention on outbound links, "ChatGPT automatically includes the UTM parameter utm_source=chatgpt.com in referral URLs." No impression, citation or appearance reporting is offered. Anthropic's publisher-facing documentation covers crawling and blocking only, and Perplexity's covers crawler control and IP verification only. Neither offers publisher analytics of any kind as of 20 August 2026. https://help.openai.com/en/articles/12627856-publishers-and-developers-faq
  • Cloudflare, "From Googlebot to GPTBot: the crawl-to-refer ratio on Radar," 1 July 2025. "However, traffic referred by Claude's native app does not include a Referer: header, and we believe that the same holds true for traffic generated from other native apps as well. As such, because the referral counts only include traffic from the Web-based tools from these providers, these calculations may overstate the respective ratios, but it is unclear by how much." / For the week of 19 to 26 June 2025: "the ratios range from Anthropic's 70,900:1 down to Mistral's 0.1:1." / Method: ratios are computed by dividing HTML requests from a platform's crawler user agents by HTML requests carrying that platform's hostname in the Referer header. / "referral traffic coming from Google's ASN (AS15169) is specifically excluded from analysis here" because of prefetching driven by speculation rules. https://blog.cloudflare.com/ai-search-crawl-refer-ratio-on-radar/
  • Cloudflare, "Verified bots with cryptography," 1 July 2025. On why user-agent counting alone is unreliable: "Existing identification methods rely on a combination of IP address range (which may be shared by other services, or change over time) and user-agent header (easily spoofable)." https://blog.cloudflare.com/verified-bots-with-cryptography/
  • Ahrefs, Brand Radar methodology, updated 26 February 2026. "All prompts run through the free, publicly available web interfaces of ChatGPT, Gemini, Perplexity, Copilot, and other supported platforms to reflect typical user experiences." / "Estimated Impressions weight mentions by Google search volume to model potential exposure. This is a modeling choice, not a measured relationship: we don't claim a validated link between Google search volume and how often a query is asked inside an AI tool." / "Metrics are directional indicators, not exact traffic counts - best understood as modeled visibility signals, and not performance metrics." / "LLMs occasionally generate hallucinated or malformed links. We do not filter out hallucinated or malformed links, as they reflect real model output." Monthly query volumes are disclosed per platform. The prompt set itself is not published. Ahrefs sells the product described. https://ahrefs.com/blog/brand-radar-methodology/
  • Semrush, AI visibility data help page, retrieved 20 August 2026. "Prompt responses are captured from real requests and not via any APIs of LLMs." / "AI search and LLM responses are fast-changing and highly personalized, which means no platform can provide exact numbers on visibility." / "We source billions of real prompts from AI search clickstream data and Google's keyword dataset for AI Overviews." Corpus reported as over 317 million prompts and responses across ChatGPT, Gemini, AI Overviews and AI Mode, across 117 regional databases. The prompt set itself is not published. Semrush sells the product described. https://www.semrush.com/kb/1607-semrush-ai-visibility-data
  • Profound, answer engine insights documentation, and Peec AI prompt setup documentation, retrieved 20 August 2026. Profound describes sending prompts to answer engines daily without disclosing the set or the runs per prompt, and separately describes an index built on licensed panel conversations. Peec AI runs customer-supplied prompts on a 24-hour cycle without disclosing runs per prompt. For Scrunch, Athena, Conductor and Similarweb's AI products, no public methodology page stating prompt set, sample size, or whether queries run against the live product or an API could be located on 20 August 2026. https://help.tryprofound.com/articles/3443229936-answer-engine-insights-overview and https://docs.peec.ai/setting-up-your-prompts
  • Kirsten et al., as reported in the Martinez survey above: 4,706 queries across several surfaces in the United States and Germany, finding two-month page overlap of 18% for AI Overviews against 45% for organic Google, and 9 to 28% of decisions changing on repeated runs where temperature could be set to zero. Findings of ACL 2026. This chapter carries the figure at one remove and has not read the primary. https://aclanthology.org/2026.findings-acl.526/
  • The "AI referral traffic is about 1% of visits" figure that circulates widely could not be traced to any primary vendor publication on 20 August 2026. It is therefore not used in this chapter as a number, only as an example of an unsourced claim.
Now taking new clients · limited spots

Reading it is one thing. Running it is another.

We build and operate this system for a small number of clients each quarter. Book a session and we'll audit where you currently sit in the pipeline.

See case studies Book Strategy Session →