Library/ GEO Playbook 2027/ Chapter 07
Chapter 7 of 12

Structured Data for AI Search: What the Evidence Shows

Structured data was supposed to be how machines read your site. Then three new protocols shipped in ten months. Not one adopted it. Nobody is asking you for it Start with what the engines say. Short, and unanimous. Google’s own guide is not hedged: “Structured data isn’t required for generative AI search, and there’s no […]

6 min read Updated Aug 2026 Part II · The Page

Structured data was supposed to be how machines read your site. Then three new protocols shipped in ten months. Not one adopted it.

Nobody is asking you for it

Start with what the engines say. Short, and unanimous.

Google's own guide is not hedged: "Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add. However, it's a good idea to continue using it as part of your overall SEO strategy, as it helps with being eligible for rich results on Google Search."

Bing's grounding guidance names four things: facts stated directly, consistent entity naming, one topic per URL, key information near the top. Not markup.

OpenAI has published nothing across its commerce documentation. Microsoft has published something.

A Bing post from May 2025 pairs schema.org Product markup with IndexNow, and says it helps engines surface content "in search results, shopping experiences, and AI-driven assistants."

Read what that is. A recommendation, with no result behind it.

Google co-founded schema.org in 2011, with Bing and Yahoo.

Fifteen years on, it will not claim the vocabulary does anything for the product it now sells.

No engine has ever published a controlled result showing that adding structured data changed an AI answer.

Has anyone shown you one? Has anyone quoted you one? Not one. In no direction. Ever.

Every one had opportunity and incentive. None took it.

Read that absence for what it is made of. It does not mean markup is worthless. It means the AI story attached to it has no published result behind it. None, from anyone.

Count the doors

Between September 2025 and June 2026, three machine-readable protocols were published. Three publishers: Google, OpenAI with Stripe, and Chrome, which is Google again. Two older pipelines already ran.

Every one is a decision about what a program should read. How many have you read? Which is your budget aimed at?

Table 7.1 · What the machine interfaces actually consumeFour protocols, two guides, two older pipelines, one proposal
InterfaceWho, and whenWhat it readsSchema.org?
Universal Commerce ProtocolGoogle, January 2026A JSON profile at a well-known URLNo
Agentic Commerce ProtocolOpenAI with Stripe, September 2025A delimited flat file, plus a REST checkout APINo
WebMCPChrome, May 2026JavaScript tool registration and JSON SchemaNo
Agent-ready toolkitChrome, June 2026The accessibility tree, layout stability, WebMCPNo
Build agent-friendly websitesGoogle, April 2026The DOM, semantic HTML, visual affordancesNo
Merchant Center crawlGoogle, years oldProduct and Offer markup, converted to a feedYes
Rich resultsGoogle, years old25 documented featuresYes
NLWebMicrosoft, May 2025Schema.org, returned as JSONYes
llms.txtCommunity proposal, September 2024No engine or protocol here reads itNo
Read the dates column before the last one. Three of the four protocols were published from September 2025 onward, and none adopted schema.org. The fourth, NLWeb, did adopt it, and it is four months older than that window and has had no successor. The two older schema-consuming pipelines both predate the products this book is about, and one of them turns your markup into a feed rather than reading it as markup. The two advice documents, one from Google and one from Chrome, are not interfaces at all, and they are here because neither mentions markup. The llms.txt row is here because somebody will ask: Chapters 1 and 5 already closed it, and Google says Search ignores it.

OpenAI's commerce feed has its own field names. It will also ingest a Google-compatible feed "without renaming its columns to OpenAI field names." So the Merchant Center field list travels. The markup does not.

We searched every page of that documentation for schema.org, JSON-LD, microdata and structured data. Zero occurrences.

Then the advice. Google's AI optimization guide names three ways an agent reads you: "analyzing visual renderings (like screenshots), inspecting the DOM structure, and interpreting the accessibility tree." Structured data is not one of them.

The page it links to, web.dev's "Build agent-friendly websites," never mentions structured data. Its advice: real buttons, labeled form fields, click targets over eight square pixels. Good advice. None of it markup.

Chrome's agent-ready toolkit is the other advice document. Same silence.

When the AI guide turns to products, it names Merchant Center feeds and Business Profiles. Then one line about the commerce protocols: "Protocols like Universal Commerce Protocol (UCP) are emerging that will allow Search agents to do more." Emerging. Not markup.

Google is pointing merchants at UCP. Not at Product markup.

Which settles the hedge Chapter 2 left here, and narrows it. What the protocols read is the feed and the endpoint. Your markup sits upstream of that, and only there.

So implement UCP or ACP? Only if you sell online, and the Merchant Center feed comes first. You probably owe one anyway.

Queue UCP rather than leading. Google published a six-step merchant path, and step one ends at a waitlist. Your integration "must be approved by Google" before you go live on AI Mode or Gemini.

The manifest is half a day. The gate after it is not yours to open. Neither protocol is a marketing task.

One thing does travel from your markup into that world. Merchant Center will build a feed from your Product and Offer markup. OpenAI takes that format.

So your markup is already most of an agentic commerce feed. Not because anybody reads it as markup. Because it is a clean source of the fields a feed wants: price, currency, availability, condition, GTIN.

One row breaks the pattern, and we put it there ourselves. NLWeb is Microsoft, May 2025: a protocol for agent access to a site. Its documentation: "The response returned uses Schema.org."

So a major engine did build a schema-consuming protocol.

Once, and four months before the window this chapter counts.

Not one release since went back to it. Including Microsoft.

One objection to our own table, and it is fair. Schema.org is a vocabulary, while the three recent protocols are transport layers. So in one sense a vocabulary is not what they declined.

Except a transport layer still names its fields. UCP wrote its own profile. ACP wrote its own field names. WebMCP took JSON Schema.

Three chances. Three vocabularies.

None of them this one.

Which is why the claim this chapter defends is the narrow one. No machine-readable protocol published since September 2025 has adopted schema.org as its vocabulary.

Take that version. It survives.

Three studies, none of them ours

One controlled study isolates JSON-LD as a variable, from March 2026. The authors are WordLift, and their product is schema markup. Read the section heading. They wrote it.

Section 4.2, heading and opening line

"H1: JSON-LD Alone Does Not Significantly Help. Adding JSON-LD structured data to HTML documents yields a small but statistically significant improvement in accuracy, though the effect size is negligible."

Their setup: a research pipeline over a closed corpus, answers scored one to five by a model judge. Not a production engine, and they say so.

Accuracy moved from 3.62 to 3.89 on that five-point scale. Effect size: 0.18. Completeness did not survive correction for multiple comparisons.

In the discussion they gloss it. Schema.org markup "remains valuable for search engines with dedicated parsers (Google, Bing), but it provides no measurable benefit in RAG-based systems that treat pages as flat text."

Keep both halves, then notice a schema vendor published the second one. When did yours send it to you?

Now the number the field quotes: +29.6%. That figure belongs to their "enhanced entity page," which adds natural-language summaries, breadcrumbs, linked entity navigation, agent instructions and a neural search reference on top of the markup. Six components, one number.

Here the authors do something unusual. They report an ablation nobody performed.

On removing the JSON-LD from the winning page they write only this: "we expect it would not" change performance, "since the same information is already expressed in natural language."

The people selling markup expect their winning condition to work without the markup. Read that twice.

Then the mechanism. Their pipeline ingests each page as one text field, truncated at roughly 20,000 characters before embedding. One field. One cut.

In their corpus, 88% of the JSON-LD pages ran past that limit. The block starts at median character 18,510.

Figure 7.2 · Where the markup sitsMedian position against the embedding cut
ONE PAGE, INGESTED AS ONE TEXT FIELD Body text, headings, navigation, boilerplate 20,000 · CUT JSON-LD starts here CHAR 18,510 88% of the JSON-LD pages ran past the limit. The markup starts in the last 1,500 characters.
This is one pipeline, not every pipeline, and the paper says so. But the shape generalizes to anything that reads a page as flat text, which is most of what people build. Your markup sits at the bottom of the document by convention, and in a system with a token budget the bottom of the document is where the budget runs out. Nothing is filtering your markup. It is arriving late.

Which gives you the cheapest move in this chapter. Put your JSON-LD block in the head, not the bottom.

Most plugins inject it wherever is convenient, and in our experience that is often the footer.

Start at 18,000 against a 20,000-character budget and a flat-text reader sees almost none of it. Start at 900 and it sees all of it.

No parser changes, no vendor, one template edit. Step 7 gives you the command that finds it.

One caveat: this is inference, not a published finding.

It holds only where the extractor keeps the script block and document order. Plenty drop script contents entirely. Chapter 3's ChatGPT measurement is that case.

Where the block is dropped, head placement buys you nothing. It costs you nothing either. Which is why we print it anyway.

Structured data has a dedicated parser in the systems that predate language models. No documented parser in any retrieval system anybody has published details of.

The same paper is plain. Google and Bing "have evolved specialized crawl-time parsers" that pull the block out as a separate signal. There, "the structured data is never flattened into a single text embedding."

One caveat on the parser side, and it cuts against us. Bing reads structured data, and Microsoft says so itself. It is "one of the clues that Bing uses."

Bing also grounds Copilot. And Google's own guide says AI Overviews and AI Mode retrieve from the Search index. Which is the index that parser feeds.

So two live paths exist where a parsed block sits upstream of an AI answer.

Neither vendor says what the parsed block does once it is there. Nobody outside has measured it either.

Two more studies, neither of them ours

Ahrefs published the before-and-after in May 2026. They sell SEO tooling, nothing in structured data. Worth copying: 1,885 pages that added JSON-LD between August 2025 and March 2026, against 4,000 matched controls. Thirty days either side.

Three surfaces. Google AI Overviews came back at minus 4.6%: statistically significant. AI Mode at plus 2.4%. ChatGPT at plus 2.2%. Their sentence: "adding schema produced no major uplift in citations on any platform."

They are careful about the negative one, and so are we. They call it "small enough that we can't definitively pin it on schema." Now their limitations. That is the interesting part.

The sample was pages already heavily cited, over a hundred citations before treatment. So it does not test a page nobody has found. All schema types were pooled.

And they looked only at markup in the page HTML. Not markup injected by JavaScript. They add that crawlers "appear to treat the two differently."

Thirty days is a short window. They cannot separate schema from anything else changing. But take it for what it is. Disinterested party, large sample, matched control.

Then a third leg, after WordLift and Ahrefs. Fischman, February 2026, published by a vendor selling GEO services.

The design: 730 AI citations from ChatGPT and Gemini across 75 commercial queries, 1,006 unique pages, with domain authority as a covariate.

Observational, not experimental. Nobody added markup to anything. Every page in it already was what it was.

The answer was null: odds ratio 0.678, p of 0.296.

Read what they found instead, worth more than the null. Rank position dominated, cutting citation odds by roughly 24% for every place you fall. Position, not markup.

Hold that loosely too. An observational rank coefficient cannot separate rank from whatever produced the rank. Nobody can.

What would your own thirty days look like? Chapter 11 builds that.

The reconciliation Chapter 3 asked for, and where your money is

Chapter 3 left this chapter a bill. A 2026 measurement found ChatGPT converting fetched pages into a stored copy, with the JSON-LD taken out. A chapter was owed on whether markup mattered.

Here is the second fact. Where the block is not removed, it is arithmetic. The markup is there, and the window closes first.

Removed, or late.

Either way it does not reach the model.

Then searchVIU, where a value existing only in markup was found by none of five systems. Perplexity was one, and Chapter 3 hedged that it might be the exception.

On this evidence it is not. One test, so hold it loosely.

Which surface is your money in?

One question, asked of every commercial page you own. Does the money arrive through a surface with a parser, or one without?

Parsers sit in Google results, Bing results and Merchant Center. Anything rendering a rich result.

Nothing else has one.

Now your numbers, not this chapter's. What share of revenue lands through classic search?

If most of it still does, your markup is doing a job. Add the two free properties later in this chapter and nothing that costs money. Then let Artifact 7.5 tell you what to remove.

If you are watching that share fall, markup does not follow the traffic. Chapters 4, 5 and 6 were the answer to that. This chapter is its boundary.

So the answer splits by what you sell. Which one are you?

Five readers who get a different answer

First, the local business, or anyone running locations. Keep your LocalBusiness markup and stop thinking about it.

It renders a panel, and that panel appears in classic results and in Maps. Both parsed, both where local revenue lands. Your hours, address, phone and geo are read there.

Locations change the arithmetic, not the verdict. Every location page runs one template. So your audit is one page and your fix is one deploy.

Which inverts the rest of this chapter. Here the surface consuming your markup is the surface paying you.

Second, the reader this book is mostly written for. You sell B2B software.

Look down the current 25. Products, locations and recipes are not yours. Most of the rest describe content you do not publish.

What is realistically left is Organization, WebSite and Breadcrumbs, which you already ship, plus Article on the blog, and Video or JobPosting if you host either yourself.

Strip the conditionals and your estate is a logo, a site name and a breadcrumb trail. All real, all small, all already deployed.

So your annual schema spend is nothing, and your project is one afternoon of Artifact 7.5. Then never again until Google adds a feature you can earn.

Third, and the only reader here told to extend markup you already ship. You sell products online.

Your Product and Offer markup is a feed. Merchant Center builds one from it. OpenAI accepts that format. The pipeline is binary.

So step 4 is most of your audit. Diff your markup against the Merchant Center product data specification. Fill the gaps.

Then know what you are buying. Google routes merchants at feeds and protocols, not at markup. So the chain runs markup to feed to answer, with a hop in the middle you do not control.

Documented at every link.

Measured at none.

One more free thing while you are in there. Bing pairs structured data with IndexNow. Their line is worth having: IndexNow tells engines "that something has changed, while structured data tells them what has changed."

If your prices move, submitting on change is one endpoint and costs nothing.

Fourth, the site that ships nothing. New build, or a theme that emits no markup.

You have no estate to audit, so the question inverts. What is the floor?

Organization on the home page. WebSite on the home page too, because a site name preference is live in every language. BreadcrumbList from the template. Product and Offer if you sell things.

That is the whole floor. Nothing else goes on until a feature exists you can earn.

Fifth, the reader who does not own the generator. Wix, Squarespace, a plugin you cannot fork.

Most of this chapter's repair is a template edit and you do not have one. So your version is shorter.

Run the audit to learn what you ship, stop paying anyone to extend it, and take whatever free moves your platform allows.

If you also sell products, the third verdict still stands. The feed fields are worth filling even on a platform you cannot template-edit. The app or the crawl carries them.

What markup actually buys

Three studies, no engine result. All of which would be a reason to stop, except structured data does pay. Just not where it is sold.

One qualifier on that Merchant Center row. If you already push a feed from a PIM, this buys you nothing.

The markup path is for merchants whose product page is the only place the data lives. What it saves is a pipeline, not a listing.

Table 7.3 · What markup is documented to buyNone of it is an AI answer
What it buysWhat it needsHow you verify itBinary?
Rich result eligibilityOne of the 25 documented features, valid and matching visible contentRich Results Test, then the Search Console enhancement reportEligible, not guaranteed
A product feed without a feedProduct and Offer markup with price, currency, availability, conditionMerchant Center builds the feed, or it does notYes. It works or it fails
Effects with nothing to look atnaics and iso6523Code on your Organization markupYou cannot. Nothing renders, so Google's documentation is the whole evidenceNo
Rows one and two you can verify yourself this week. Row three you cannot, and it is the only place in this chapter where the standard of evidence drops to a vendor's own documentation. We print it because both properties in it are free and take ten minutes. Hold it to a lower confidence than everything above it. The second row is the strongest case for structured data in 2026 and almost nobody in this field talks about it. Price and availability stay in sync from the same source your customers see, which makes it a data integration with a pass or fail outcome. That is the opposite of everything else in this chapter. If you sell products online, this row justifies the work on its own, and it has nothing to do with being cited by anything.

Rich results are the older case. How short is that list now?

HowTo went in September 2023. The sitelinks search box in 2024. Practice problems in 2025. FAQ stopped appearing on 7 May 2026. The documentation came down 15 June.

Notice how that one was announced. No blog post: a changelog entry, and the old URL redirects to it. Which is why the advice outlived the feature.

Now plot the four. One loss a year, four years running. Nothing added for a generative surface.

So name the survivors. Step 2 of your audit needs something to check against.

The features still rendering are the commercial and the calendar ones. Product and merchant listings. Review snippets. Job postings. Events. Recipes. Videos. Breadcrumbs.

Notice what all but one share. What do they describe? A thing with a price, a date or a duration.

Breadcrumbs is the exception. It renders a path rather than a fact.

Which is the rule underneath the gallery. Google renders markup carrying a fact a search result can display, and stops when it does not. FAQPage carried an answer. Answers are what the generative surfaces now do themselves.

So read the current 25 as a ceiling that drops.

Publishers get one more item, too small for Table 7.3. Google states plainly that there is "no markup requirement to be eligible for Google News features like Top stories." What Article buys is better title text, images and dates. Real, and smaller than a feed.

One publisher decision matters more: the paywall. Running a subscription wall? Then set isAccessibleForFree to false and add a hasPart block naming the walled section by CSS selector. Two properties, one decision.

That shows Google the full text without being treated as cloaking.

It changes how a surface treats your content, not how it renders it.

Which puts it in row one, not row three. Search Console carries a Subscribed content report. You can see whether it worked.

Traced to source

Chapter 5 traced five writing prescriptions. Chapter 6 traced five entity claims. Same method, third time.

Table 7.4 · Five schema claims, tracedOrigin document, and whether it holds
The claimWhere it comes fromDoes it hold?
"Microsoft confirmed schema helps LLMs"A conference attendee's LinkedIn post about a stage remark. The "confirmation" is a seven-word reply to that postNo Microsoft document says it changes an answer
"Use @id to build a knowledge graph"Google documents @id as in-page plumbing: policy nodes, loyalty tiers, and linking a recipe and a video "on the page"In-page plumbing, three cross-page links, one feed spec you are not in
"sameAs builds your entity"Google calls it "a page on another website with additional information about your organization," giving a social profile as the exampleAdditional information, not identity
"Add FAQPage schema for AI answers"A rich result that stopped appearing on 7 May 2026The feature does not exist
"Speakable schema optimizes you for voice AI"A Google Assistant feature for reading news aloud on smart speakers, released in July 2018 and still marked betaUS, English, news publishers
Row four is one grep, and the fastest test you can run on any source. FAQ rich results ceased in May 2026 and the documentation came down in June, so a 2026 guide still recommending FAQPage markup for AI visibility has not read a primary source this year. It tells you whether to keep reading. Row five is the subtler tell: speakable is still in the gallery, still in beta after eight years, and presence in the gallery is not evidence of relevance.

The first row is worth walking. A Bing product manager spoke at a conference in Munich, March 2025.

An attendee, David Mihm, wrote a LinkedIn post summarizing it: schema markup helps Microsoft's language models understand your content. The product manager replied. His entire reply: "Thanks, please to help all of you."

Two trade publications called it confirmation. There is no Microsoft document behind it.

The nearest thing is that May 2025 Bing post, co-authored by the same product manager. It shows nothing.

And the claim moved in transmission. Mihm's sentence became the field's line that schema is how you talk to AI. One is a remark about comprehension. The other, a claim about a channel.

How many claims in your last schema proposal would survive the walk?

The threads Chapter 6 left here

Two properties exist to say this is the same entity as that one, and Chapter 6 left them here. Start with sameAs, the one every entity retainer bills for.

Google's Organization documentation lists twenty-three properties. Seven are identifiers: DUNS, GLN, ISO 6523, LEI, NAICS, tax ID, VAT ID. Five are numbers somebody issued you. One wraps a number you already have. One you pick from a table.

Google's own gloss for sameAs: "The URL of a page on another website with additional information about your organization, if applicable." Additional information. Nobody issues it, and nobody checks it.

The page never uses the word entity, except inside "Legal Entity Identifier." Knowledge Graph does not appear at all, and no property is marked required.

No threshold to reach.

Nothing to buy.

Now the better thing Chapter 6 handed over. It found Google using the word disambiguate in one place. Its Organization documentation.

What it could not do was price the two codes it names. This chapter can.

Some properties, it says, are used "behind the scenes to disambiguate your organization from other organizations (like iso6523 and naics)."

That is the page's only claim about a function you can act on. Two other properties carry non-rendering language: url "helps Google uniquely identify your organization," and vatID is "an important trust signal for users," who can look you up in a public VAT registry. Neither is something you choose.

Google's own split is two-way. Some properties work behind the scenes, and others "can influence visual elements in Search results."

What the rest do it never says. And it names nothing an entity retainer has ever billed for.

Neither is a registration either. NAICS is an industry code you pick yourself from a public table. ISO 6523 wraps a number somebody already gave you.

Both free.

Both ten minutes.

So the ruling on all seven. Copy across every registered identifier your finance system holds. Then fill in naics and iso6523Code, the two Google gives a function and no rendering. The property carries the Code suffix even though Google's prose drops it. Do not obtain a number you do not have.

One limit, held to the same standard this chapter holds everybody else to. The documentation locates the disambiguation "in search results" and stops there.

No feature named. No measurement. Nothing you can look at.

We print it because it costs nothing. Not because anybody has shown it does anything.

That is Chapter 6's finding, cashed.

Then @id, which Chapter 6 also left open. Google documents it in several places, and every one is plumbing.

Joining a Product to its shipping and return policy nodes, so you write the policy once. Linking loyalty tiers. Linking a recipe and a video "on the page," so the video can appear inside a Recipe rich result.

Some of it does cross a page boundary. Merchant listing uses @id to point an Offer at your global return policy, your global shipping policy and your member tiers. Each by absolute URL.

Three cross-page uses. Two point at a policy page you already publish. The third at a members page.

So the honest reading: @id stops you repeating yourself, inside a page and across three named pages you already run. One feed spec you are not in requires it to be globally unique. Nothing documents it as a way to build a site-wide graph.

Which leaves the residue of both. Organization markup helps Google understand your administrative details and render your logo. Two codes help it tell you apart from somebody else. Real, modest, checkable. Not entity construction. Does your invoice say otherwise?

The one rule, and the thing that voids it

One rule sits underneath all of this, from Google's own quality guidelines. Every fact worth marking up has to be visible text on your page.

Two exceptions, both already in this chapter. Identifiers nobody renders, and a paywall block Google documents as the way to show it what a reader cannot see.

Everything else obeys.

That is not a style preference. It costs you rich result eligibility. Google's structured data policies carry a manual action for markup that does not match visible content. A real penalty.

The rule also makes markup work where nothing parses it. If the price sits in your JSON-LD and in your paragraph, the flat-text pipeline finds the paragraph. If only in the JSON-LD, nothing finds it.

So the policy rule and the machine-readability rule are the same rule. Which almost never happens.

Artifact 7.5 · The markup decisionSeven questions, one afternoon
  1. List the types you currently ship. One URL per template, twelve templates, chosen by revenue rather than page count. If you have forty, take the twelve carrying the money, because this audit finds template defects and twelve will find them. Not what the plugin claims: paste each URL into validator.schema.org, which reports every type on the page rather than only the ones Google supports, and write down what comes back. The Rich Results Test cannot do this step: it is silent about types Google no longer supports, which is exactly the junk you are looking for. Most sites are shipping types nobody remembers adding.
  2. Cross off anything Google no longer documents. Start with the gallery, which lists 25 features and contains neither FAQPage nor HowTo. Then check the ecommerce pages that sit outside the gallery: merchant listing, product variants, loyalty program, return policy. If Google has no page for the type at all, it is maintenance you are paying for and receiving nothing back.
  3. For each surviving type, name what it earns. For most types that is a rich result: run the Rich Results Test, then open the Search Console enhancement report for that type, and no valid items means it is not earning its place. Nine of the 25 have no enhancement report at all: Article, Carousel, Course list, Employer aggregate rating, Local business, Movie, Organization, Software app and Speakable. An absent report is not the same as an absent result. So the rule is: no report and no documented function means the type fails. Four types have a documented function anyway. Organization renders your logo. LocalBusiness renders the panel in Search and Maps. Article states title text, image and date preferences in Search and Google News. And WebSite, which is not in the gallery at all, states your site name preference. Score all four on the function. While you are in Search Console, open the manual actions report once: a structured data action costs you eligibility and you will not be told any other way.
  4. If you sell products, check Merchant Center separately. This is the one row that is binary. If your products already arrive by feed, app or PIM, this row is bought by other means: record it, run the JavaScript check below anyway, and skip the rest of this step. Otherwise open Merchant Center, then Products, then the diagnostics view: either the website crawl is building items from your Product and Offer markup or it is reporting them as missing required attributes. Then run step 7 on a product URL in the same pass, because a JavaScript-injected block is the silent failure here.
  5. Then run the visible-text check. Ten marked-up values at random, drawn across the whole sample rather than from one page. Sample only value-bearing properties, meaning anything a customer could read as a fact: price, rating, author, date, availability. Skip URLs, identifiers, image references and enumerations, which are not what the visible-content rule is about. For each one, find it in the visible text of the same page. Any value you cannot find is both a manual action risk and a fact no flat-text pipeline will ever see. If fewer than ten value-bearing properties exist across the sample, score all of them and report the denominator you used. This step scores separately, as a fraction out of ten, and the pass mark is ten.
  6. Cancel anything that survives none of the above. A type with no rich result, no Merchant Center role, and no visible-text backing is doing nothing measurable in any system anybody has tested. Deleting it is the deliverable. If a plugin or a hosted platform emits it and you cannot remove it without disabling the plugin, deletion is not available and not worth a migration: record it, stop paying to extend it, and move on.
  7. Then do the three free things while the template is open. Run curl -s YOURURL | grep -bo 'application/ld+json' on each of the twelve. No match at all is the finding: the block is JavaScript-injected, Merchant Center will not take it, and Anthropic's fetch tool cannot see it, so move it server-side. The number before the match is a byte offset. Any block starting past byte 15,000 moves into the head, which is our margin against the roughly 20,000-character cut the WordLift pipeline used rather than a published threshold. Then add naics and iso6523Code to your Organization markup, because they are the two properties Google names as working behind the scenes, and neither costs anything. Last, curl one known duplicate URL, parameterized or paginated, and confirm the block is there. Absent means your generator emits on canonicals only, which Google's duplicates guidance recommends against.
Sample, scorer, cadence. Twelve templates, or all of them if you have fewer than twelve. The person who owns the template runs it, not the person who sold you the schema. Re-run every six months if your estate changes. If your verdict was B2B software or freeze, run it once and re-run only when a changelog entry removes a feature. Subscribe to the Search Central changelog page, because that is where FAQ was killed and there was no blog post. The trigger matters more than the calendar.
What it returns. One point for every distinct type that clears step 3 or step 4. Distinct types, not instances: a plugin graph emitting four types across twelve templates is four. Divide by the count of distinct types you found in step 1, before any crossing off, and read the percentage. If step 1 returned nothing, stop: you have no estate, and your job is the greenfield floor rather than this audit. Step 5 is reported alongside it as a separate fraction out of ten, because it scores values rather than types. Its pass mark is ten, and that one is not ours: the visible-text rule is a policy requirement, so the tenth miss is no worse than the first. Anything under ten is a template defect rather than a quality gradient. On the type percentage, the floor is fifty.
Why fifty. Below half, your estate has stopped being a deployment somebody decided on and started being plugin output. That is a different repair. You change the generator once rather than editing types one at a time, and it is cheaper. One exception, and it is the common one: an SEO plugin emits a fixed graph of five or six types whether they earn anything or not. If your misses are all plugin boilerplate you cannot remove without disabling the plugin, exclude them from the denominator and score what you control. Report both numbers. The unexcluded percentage is what your estate is. The excluded one is what your decisions are, and if the gap is more than twenty points the gap is the finding: change the generator, not the types. If exclusion leaves fewer than three types, the percentage does not apply at all and your verdict is freeze plus steps 5 and 7. Fifty is ours, and it is a floor rather than a target: raise it after you have run the audit twice.
What this does not measure. Nothing here tells you whether markup affects an AI answer, because nobody has published a way to tell. It measures what your markup is documented to buy and whether you are getting it. Two of the seven steps can only ever return "cancel this", which is the honest shape of the evidence rather than a rhetorical choice.
Five instruments, none of them licensed, and one command. Validator.schema.org for what you actually ship, the Rich Results Test for eligibility, the Search Console enhancement and manual action reports for whether eligibility became anything, Merchant Center for the one binary case, the Search Central changelog for when a feature dies, and curl for position and for whether the block survives without JavaScript. Nothing here needs a crawl of your own. The scoring floor exists so the audit ends in a decision rather than a spreadsheet. It, the twelve templates and the 15,000-byte trigger in step 7 are the three numbers the scoring rests on, all ours, each printed with its reasoning. The sample sizes and the six-month cadence are ours too, and nothing turns on them.

Price the cost side first

Nobody prices the cost side, so price it before you defend it.

Three lines, all pullable this week.

The schema line on your last three invoices, agency or plugin subscription. Then engineering hours per template change, times the templates you touched last year.

Then the one nobody books: every price, availability and date exists twice, and somebody keeps the copies in step. Second syntax, second system. Usually a plugin.

Google requires the copies to agree. Merchant Center puts it hardest: structured data "must match the values that are shown to the customer." And the duplicates guidance doubles the surface. Put the same structured data "on all page duplicates, not just on the canonical page."

Which is a template-generator property, not a page one. If your markup is server-side it is already handled. If it arrives through a tag manager it is not.

So the failure mode is not that your markup is wrong. It is subtler. When did your markup last describe this month's price?

Vendors overstate the consequence. A structured data manual action costs you rich result eligibility and nothing else. Anyone telling you schema errors will tank your rankings is contradicting the documentation, in the sentence that describes the penalty.

The thing Merchant Center will not take

One more cost nobody checks. Google Search processes markup injected by JavaScript. Merchant Center does not: the markup "can't be generated with JavaScript after the page has loaded."

So the one binary pipeline here rejects it. Anthropic's web fetch tool cannot see it either. It "currently does not support websites dynamically rendered with JavaScript."

Another free move, then. If the block only appears after render, move it into your server-side template. Step 7 finds out in one command.

What a crawler can execute belongs to Chapter 10.

The installed base is mostly the smallest types

One verdict before the numbers, for anyone on an inherited microdata estate. Do not migrate it.

Microdata is still parsed, and migration is a template rewrite bought with the money this chapter tells you to stop spending. Convert on the next template touch, or never.

JSON-LD is on 41% of pages, up from 34% two years earlier.

That is the 2024 Web Almanac.

Now the four commonest on mobile. WebSite: 12.73%. Organization: 7.16%. BreadcrumbList: 5.66%.

LocalBusiness: 3.97%.

WebSite powered the sitelinks search box, retired in November 2024. It is also how you state a site name preference, live in every language on both devices. Organization renders a logo. BreadcrumbList renders a trail. LocalBusiness renders a panel.

Each buys one small rendering. Each is real. Then Product, the binary row, at 0.77%. So the web deployed the markup with the smallest payoff. At sixteen times the rate of the largest.

Both shares are measured over all pages, which flatters WebSite. Product only makes sense on a merchant site.

Most companies never chose to deploy schema. A plugin did, or an agency did, four years ago.

So your question is not what to add. What do you keep?

Artifact 7.5 answers that. Keep what earns a report or a documented function. Keep whatever Merchant Center consumes. Delete the rest.

Quote Google's line at whoever objects. Unused structured data "has no visible effects in Google Search."

It is not hurting your rankings. It costs you every time somebody changes a template. One caution before you start deleting. Markup is a template property, not a page property. So audit one page per template, and change it once.

If the estate is genuinely large, the verdict changes shape. Forty templates, a vendor contract, tag-manager deployment.

There, deletion is not the deliverable. Freezing is.

Stop paying for expansion this quarter. Run Artifact 7.5 once, across templates not pages. Delete only on the next scheduled template touch. A deploy you were making anyway is free. One you scheduled for this is not.

A different answer from the one a ten-template company gets. The difference is the cost of the deploy. Not the value of the markup.

The objections this chapter has to answer

One thing to fix before any of them. This chapter is not telling you to cancel markup, only the AI story attached to it, and Artifact 7.5 keeps whatever earns a rich result or a feed.

The first objection, in two halves.

"Absence of evidence is not evidence of absence. Web Data Commons pulls schema.org out of Common Crawl and publishes it as tens of billions of triples, so my markup demonstrably sits in a public dataset model builders use. Nobody has shown it does nothing at the training stage, and you cannot see that stage. And engines have every reason not to tell you. Publishing what moves an answer is publishing a ranking factor."

Both halves are fair. Silence is consistent with markup working, and equally consistent with markup doing nothing.

It cannot be evidence for one. What it settles is who pays for the uncertainty, and right now that is you.

Then weigh what the silence is not made of. The vendor with the most to gain published a study whose section heading says its product does not help.

Not proof.

The strongest thing a schema seller has printed against itself.

The second, and it is the one every CFO raises.

"The markup is deployed. It costs me nothing incremental. Google co-founded schema.org, and if some future protocol adopts it I am ready for free. Your spine is a snapshot of June 2026."

The snapshot point is correct. This chapter counted the doors on one day.

The option breaks on carrying cost, which you priced earlier. Two copies that must agree, forever, against a gallery losing a feature a year.

You are carrying a call option whose underlying keeps shrinking.

A third, built out of our own third leg. "Rank dominates citation odds, and rich results move click-through. So markup buys rank, and rank buys citation."

It breaks at the middle link. Eligibility is not rank, and nobody has published a rich-result-to-rank effect, so the chain has a first step and a last step and nothing in between.

A fourth, and the only one that names a mechanism rather than a result.

"Your truncation evidence is one pipeline's implementation choice. In a grounding step that passes a hundred thousand tokens of raw HTML to a model, my JSON-LD is inside the window. It is also the most token-efficient description of my facts that exists."

A real mechanism, and we cannot rule it out. Nobody has published a grounding step in enough detail to test it.

Two things constrain it. First the visible-text rule, which Google requires anyway. If every marked-up fact is also in your prose, the model has it regardless.

Then searchVIU, from earlier in this chapter. Truncation does not explain that miss. The JSON-LD price sat earlier in the document than the visible price the same systems did return.

One page, one test, and the only one anybody has run on this question.

A fifth, and it is built from this chapter's own sources. It arrives with a preamble worth granting first: the vocabulary really is everywhere. Pinterest documents Rich Pins on schema.org, and every enterprise retrieval stack with an off-the-shelf extractor reads it, because it is the only vocabulary there is.

Granted, and none of those is a surface citing you. Then the objection proper.

"Merchant Center builds a feed from my Product markup. Google's AI guide says feeds are how products get into AI responses. OpenAI takes that feed format. You documented every link. That is markup reaching an AI answer."

Every link is documented, and we printed all three. So concede it: for a merchant, a path exists.

Two things bound it. The hop in the middle is a feed, which is why the instruction here is fill the feed fields rather than expand your schema. And nobody has published what the feed changes in an answer either.

A sixth, from the reader this chapter is hardest on. "I signed a schema SOW three weeks ago. What do I do on Monday?"

Redirect it. The same team does the work either way. Do not cancel it.

Ask for Artifact 7.5 plus a Merchant Center feed check, not a site-wide expansion. Smaller scope, same budget, and it returns a number you can read.

Then ask for the origin document behind every claim in the proposal. If those come back as blog posts, you have learned how your retainer was priced.

A seventh, and it is the sharpest thing anyone can say about this chapter. "You refused every claim backed only by a vendor's documentation, then told me to add two properties on exactly that basis."

Correct, and Table 7.3 concedes it in the row itself. The answer is price rather than evidence: ten minutes and zero dollars is a different decision from a retainer.

So that row sits at the bottom of this chapter's confidence order. We put it there ourselves.

Which leaves the one we raise against ourselves. You say the field cannot show a result, so where is ours?

We do not have one, and nobody does. The rule here is about burden rather than evidence: the party proposing the spend produces it, which is why Artifact 7.5 runs on instruments Google hands you free.

Now the disclosure. We sell GEO services, and have billed for schema work.

We are telling you to shrink it to one afternoon. Then notice where the canceled budget goes. Into Part III, where the off-site work lives. A bigger engagement than a schema retainer.

Which is a reason to ask us the question, not a reason to disbelieve the chapter. Ask it.

What is left

A survey published in July 2026 read 45 GEO studies from November 2023 onward. Its conclusion is the right place to stop. No reviewed technique "shows a stable, longitudinal, cross-platform causal effect on organic discoverability."

Not structured data. Not anything else. So keep the markup that buys you a rich result. Keep the markup building your feed.

Make every marked-up fact visible. It is a policy requirement. And the only form a flat-text system can read.

And stop paying for the rest of it. This chapter closes Part II.

Everything since Chapter 4 has been about the part you control completely, and none of it decides whether an engine has heard of you.

Four chapters converge on one instruction. Every fact a machine needs has to be visible prose. Inside a span that survives being cut out of your page.

Chapter 4 got there through chunking. Chapter 5 through vocabulary. Chapter 6 through naming. This chapter got there through a policy rule Google wrote for a different reason. That is the closest thing to a finding this part has.

Not one of the four can tell you what a production engine did with any of it.

One conflict stays open, and pretending otherwise would be dishonest.

In the two studies that put rank and AI-specific work in the same frame, the AI-specific work does not clear its bar and rank does. Neither was built to compare them, which is itself the finding.

Chapter 4 had an escape from that. It said C-SEO Bench measures what happens after retrieval, not whether you get retrieved.

The third leg closes that escape. Fischman is observational, on live engines. Rank dominates at the retrieval end too.

We have half an answer, and it is the one from earlier: an observational rank coefficient cannot separate rank from whatever produced the rank.

That weakens the argument. It does not dispose of it. Nothing in these four chapters does.

Three debts leave this chapter unpaid, named so you can hold us to them.

Markup injected by JavaScript has an address, and it is Chapter 10.

Head placement needs a test somebody can run: the same page, published twice, block in the head and block in the footer, then ask the engines. We have not run it.

And the Bing parser path to Copilot has no address at all. Nobody outside Microsoft can settle it. So it stays open for the life of this book.

Now the thing Part II could not answer. Every audit assumes a system that has already retrieved your page. Chapter 8 owes you the step before that. Why a model names some brands and not others.

Chapter 8 is about the part you do not own. That is the debt this part hands over.

Chapter 07 · What to do with this:
  • Delete FAQPage and HowTo. Both rich results are gone, and one lost its documentation in June 2026 with no blog post. Grep any proposal for either name.
  • If you sell products, check Merchant Center first. Product markup becomes a feed and keeps price in sync. Binary, and the strongest case in this chapter.
  • Move the JSON-LD block into the head. One curl to find it, one template edit to move it. Costs nothing, and it is inference rather than a finding.
  • Fill in naics and iso6523Code. Ten minutes, free, the two properties Google names as working behind the scenes, and nobody sells them.
  • Ask any schema claim for its origin document. The loudest one in this field traces to a LinkedIn post about a conference remark.
  • Do not buy the AI story. No engine has published a controlled result, and three studies found effects too small to act on.
Sources
  • Google Search Central, "Optimizing your website for generative AI features on Google Search," updated 10 July 2026. "Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add." The same guide states that agents work by "analyzing visual renderings (like screenshots), inspecting the DOM structure, and interpreting the accessibility tree," names Merchant Center feeds and Business Profiles for getting products into AI responses, and writes that "Protocols like Universal Commerce Protocol (UCP) are emerging that will allow Search agents to do more." It links to the separate web.dev page "Build agent-friendly websites," which carries none of those three lines.
  • Google Search Central, "Structured data markup that Google Search supports," updated 15 June 2026. Twenty-five documented features in the gallery table. No faqpage and no how-to entry remains in the gallery.
  • Google Search Central, "Site names in Google Search." "To indicate your site name preference, add WebSite structured data to your home page." Site names are available in all languages where Google Search is available, on both mobile and desktop. Google Search Central Blog, "Farewell, Sitelinks Search Box," October 2024, with the feature ceasing to appear 21 November 2024.
  • Google Search Central, Article structured data. "While there's no markup requirement to be eligible for Google News features like Top stories, you can add Article to more explicitly tell Google what your content is about." The documented benefit is better title text, images and date information. Subscription and paywalled content guidance covers isAccessibleForFree and hasPart.
  • Google Search Central, Local business structured data. The panel appears when users search for businesses on Google Search or Maps.
  • Google Search Central Blog, "Changes to HowTo and FAQ rich results," 8 August 2023. Unused structured data "has no visible effects in Google Search."
  • Google Search Central, "Generate structured data with JavaScript." Google Search processes JavaScript-generated structured data. Google Merchant Center help, answer 6069143, "About structured data markup for Merchant Center," which states structured data "must match the values that are shown to the customer." Answer 7331077, "Set up structured data for Merchant Center," which states the markup "can't be generated with JavaScript after the page has loaded."
  • Google Developers, Universal Commerce Protocol implementation guides. A six-step merchant path: prepare your Merchant Center account, which ends by joining a waitlist, then set up Google Pay, publish the UCP profile, complete native checkout, choose a user identification path, sync order status. "Your integration must be approved by Google before you can go live on AI Mode in Google Search and Gemini."
  • Bing Webmaster Blog, "Introducing JSON-LD Support in Bing Webmaster Tools," 30 July 2018. "Bing works hard to understand the content of a page and one of the clues that Bing uses is structured data," and the Markup Validator "supports six markup languages, including Schema.org, HTML Microdata, Microformats, Open Graph and RDFa." Microsoft therefore documents that Bing reads structured data. No Microsoft document describes what a parsed block does to a Copilot answer.
  • Schema Markup Validator, validator.schema.org. Validates all schema.org markup on a page. The Rich Results Test at search.google.com/test/rich-results reports only the Google rich result features a page is eligible for, which is why step 1 of the artifact uses the first and step 3 uses the second.
  • Google Search Console Help, "Rich result reports," answer 7552505. Nineteen supported reports: breadcrumbs, datasets, discussion forum, education Q and As, events, hotels, image metadata, job postings, math solvers, merchant listings, practice problems, product snippets, profile pages, Q and As, recipes, review snippets, subscribed content, vacation rentals, videos. Nine gallery features have no report: Article, Carousel, Course list, Employer aggregate rating, Local business, Movie, Organization, Software app, Speakable.
  • Google Search Central, "Intro to how structured data markup works." Three supported formats: JSON-LD, recommended, plus Microdata and RDFa. Microdata is still parsed, which is why this chapter does not recommend migrating an inherited microdata estate.
  • Google Search Central, structured data documentation index. Merchant listing, product variants, loyalty program and return policy are documented ecommerce pages that do not appear as rows in the rich result gallery, which is why the artifact checks both.
  • Google Search Central, on llms.txt: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them." Restated in the changelog of 15 June 2026. The file was proposed in September 2024, and none of the other interfaces in Table 7.1 references it. Closed in Chapters 1 and 5.
  • Pinterest developer documentation, Rich Pins. Article, Product and Recipe pins, built on Open Graph and schema.org. Cited in the objections section as the clearest case of a platform that does consume the vocabulary.
  • Web Data Commons, structured data extraction from Common Crawl, webdatacommons.org. The most recent published extraction is October 2024, at roughly 74 billion quads. Cited in the objections section as the strongest form of the training-corpus argument, which this chapter concedes it cannot answer.
  • Google Search Central, general structured data guidelines and structured data policies. The duplicates recommendation, in full: "we recommend placing the same structured data on all page duplicates, not just on the canonical page." The requirement that marked-up content be visible to users, and the manual action that costs a page its rich result eligibility when markup and content disagree, and which Google states does not affect how the page ranks.
  • Google Search Central, Organization structured data. The page states that some properties are used "behind the scenes to disambiguate your organization from other organizations (like iso6523 and naics)," and that adding the markup "can help Google better understand your organization's administrative details and disambiguate your organization in search results." The property name is iso6523Code even though the prose drops the suffix. The gloss for sameAs is "The URL of a page on another website with additional information about your organization," with a social or review profile given as the example. The word entity appears only inside "Legal Entity Identifier," Knowledge Graph does not appear at all, and the page states three times that there are no required properties.
  • Google Search Central. The general structured data policies use @id to link a recipe and a video "on the page" so the video can appear as a Recipe rich result. Product variants uses it to reference shipping and return policy nodes within one page. Loyalty program uses it for member tiers within one page. Merchant listing uses absolute-URL @id references in three documented places: a global return policy page, a global shipping policy page, and a members page for loyalty tiers. The Book actions feed specification, a limited-access program, requires @id to be globally unique and stable. No general site-wide graph function is documented.
  • Google Search Central changelogs. HowTo deprecated September 2023. Practice problems announced gone November 2025, documentation removed January 2026. FAQ rich results ceased appearing 7 May 2026, documentation removed 15 June 2026, with no accompanying blog post.
  • Google Search Central, "Speakable (Article, WebPage) structured data (BETA)," updated 10 December 2025. Google Assistant reading news aloud on smart speakers. English-language publishers only. In beta since 2018.
  • Microsoft, NLWeb, announced 19 May 2025. A protocol for natural language and agent access to a website, built on MCP. Its own documentation: "It returns responses in JSON using Schema.org," and "The response returned uses Schema.org, a widely adopted vocabulary for describing web data." It is the one recent protocol from a major engine that consumes schema.org, and it predates the September 2025 window by four months.
  • Bing Webmaster Blog, "IndexNow Enables Faster and More Reliable Updates for Shopping and Ads," 19 May 2025, by Fabrice Canel and Krishna Madhavan. Recommends schema.org Product markup alongside IndexNow. "IndexNow tells search engines that something has changed, while structured data tells them what has changed." It states that structured data helps engines surface content "in search results, shopping experiences, and AI-driven assistants." It is a recommendation and reports no result.
  • Google Developers Blog, Universal Commerce Protocol, 11 January 2026, and the profile documentation it links to. A JSON manifest at a well-known URL. The strings schema.org and JSON-LD do not appear in the announcement body or in the profile documentation.
  • OpenAI with Stripe, Agentic Commerce Protocol and Instant Checkout, 29 September 2025. A UTF-8 tab or comma delimited feed plus a REST checkout API. The spec accepts a Google-compatible product feed as is, and points at the Merchant Center product data specification for field definitions. A registered Google-compatible feed can be uploaded "without renaming its columns to OpenAI field names," which is a compatibility path rather than an adoption of the Merchant Center vocabulary. The strings schema.org, JSON-LD, microdata and structured data appear zero times across every page of the specification.
  • Chrome, WebMCP documentation published 18 May 2026, origin trial announced 9 June 2026 and available from Chrome 149, and the agent-ready toolkit, June 2026. JavaScript tool registration with JSON Schema, and the accessibility tree.
  • Google web.dev, "Build agent-friendly websites," 1 April 2026. Linked from the AI optimization guide. On interactive elements: make sure they "have a visible area larger than 8 square pixels, to avoid being filtered out by visual analysis." The strings schema.org, JSON-LD, microdata and structured data appear zero times on the page.
  • Web Almanac 2024, structured data chapter, published 11 November 2024. JSON-LD on 41% of pages, up from 34% in 2022. Most common types on mobile: WebSite 12.73%, Organization 7.16%, BreadcrumbList 5.66%, LocalBusiness 3.97%, Product 0.77%. A 2025 edition exists but has no dedicated structured data chapter. Its SEO chapter does carry a structured data section, reporting JSON-LD at 43% of home pages and 39% of desktop inner pages, on a different basis from the 2024 chapter. The 2024 structured data chapter remains the comparable source for the type-level shares quoted here, and it is roughly two years old.
  • Google Merchant Center help, product data from your website and the structured data setup requirements. Markup becomes a feed through the website crawl, and the markup "can't be generated with JavaScript after the page has loaded."
  • Anthropic, web fetch tool documentation. The tool "currently does not support websites dynamically rendered with JavaScript."
  • Schema.org, announced 2 June 2011 as a joint initiative of Google, Bing and Yahoo. Yandex joined in 2011. The vocabulary Google co-founded is the vocabulary Google declines to recommend for its generative surfaces fifteen years later.
  • Bing Webmaster Guidelines, sections 15 to 18, as cited in Chapters 5 and 6. The guidelines are served as a JavaScript application, so as in Chapter 6 the four grounding items are taken from trade reporting of the February 2026 update rather than from a fetched primary rendering. Grounding guidance names content that stands on its own, consistent entity naming, one topic per URL and key information early. It does not name structured data.
  • RESONEO, August 2026, as cited in Chapter 3. Independent measurement of ChatGPT converting fetched pages to a stored copy with the JSON-LD removed.
  • Volpini, Raad, Gamba and Riccitelli, "Structured Linked Data as a Memory Layer for Agent-Orchestrated Retrieval," arXiv:2603.10700, 11 March 2026. The authors state their affiliation as WordLift, a commercial schema markup vendor. Section 4.2 is titled "H1: JSON-LD Alone Does Not Significantly Help." Accuracy 3.62 to 3.89, effect size 0.18, completeness not significant after correction. The +29.6% headline belongs to their enhanced entity page rather than to the markup, and of an ablation removing the JSON-LD from that condition the authors write only "we expect it would not" change performance. Their pipeline truncates each document at roughly 20,000 characters, 88% of the JSON-LD documents exceeded it, and the block begins at a median position of character 18,510.
  • Linehan and Guan, Ahrefs, "We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved," 11 May 2026. Difference in differences, 1,885 treated pages against 4,000 matched controls, thirty days either side. AI Overviews minus 4.6% and significant, AI Mode plus 2.4%, ChatGPT plus 2.2%. The authors note the sample was already heavily cited, that all schema types were pooled, that only JSON-LD was tested, and that they cannot separate schema from other simultaneous changes.
  • Fischman, "Does Schema Markup Predict AI Citation?", February 2026, cited in this chapter as the third leg of the field evidence. Observational across live engine results, with a citation counted as a linked source in an answer. It was published by Growth Marshal, a vendor selling GEO services. The design is 730 AI citations from ChatGPT with browsing and Gemini with grounding, across 75 commercial queries, against a control set of Google top-ten organic results, for 1,006 unique pages. Domain authority is the covariate, and the models are generalized estimating equations with query-clustered standard errors. The weakness is the design rather than the disclosure: the control set is Google's top ten rather than the same population, which the author identifies himself. The headline schema-presence result is null at odds ratio 0.678 and p of 0.296. Rank position dominates, reducing citation odds by roughly 24% per position. The Product and Review subgroup figure is a post-hoc comparison and should not be quoted as a multiplier.
  • "Optimizing Visibility in Generative Engines: A Critical Survey," arXiv:2607.14035, 15 July 2026. Reviews 45 studies from November 2023 to July 2026. No reviewed technique "shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior."
  • searchVIU, testing conducted October 2025. A single purpose-built page with prices distributed across visible HTML, JavaScript, JSON-LD, microdata and RDFa. The value present only in JSON-LD was found by none of the five systems tested. The author notes markup may still be used at earlier stages, which this test does not measure.
  • David Mihm, LinkedIn, March 2025, summarizing a Fabrice Canel presentation at SMX Munich. Canel's reply to that post, in full: "Thanks, please to help all of you." Reported as confirmation by Search Engine Roundtable and Search Engine Land, 20 March 2025. Microsoft's most recent publication announcing a Bing feature built on schema.org markup is the Bing Webmaster Blog post of 23 March 2020, "Bing adopts schema.org mark-up for Special Announcements about COVID-19." The May 2025 IndexNow post names surfaces prospectively, announces no feature and reports no result.
Now taking new clients · limited spots

Reading it is one thing. Running it is another.

We build and operate this system for a small number of clients each quarter. Book a session and we'll audit where you currently sit in the pipeline.

See case studies Book Strategy Session →