Library/ GEO Playbook 2027/ Chapter 05
Chapter 5 of 12

Writing for AI Citation: What the Research Actually Shows

Two search engines have published writing guidance for AI answers. They contradict each other. Neither has run a single experiment. The style guide nobody quotes Exactly one engine has ever told you how to write a sentence for extraction. It was not Google. Microsoft published it on 8 October 2025, under the title “Optimizing Your […]

8 min read Updated Aug 2026 Part II · The Page

Two search engines have published writing guidance for AI answers. They contradict each other. Neither has run a single experiment.

The style guide nobody quotes

Exactly one engine has ever told you how to write a sentence for extraction. It was not Google.

Microsoft published it on 8 October 2025, under the title "Optimizing Your Content for Inclusion in AI Search Answers." Four criteria, under a heading asking what makes content eligible for featured snippets. Four.

Nobody quotes them. Not one agency deck.

Microsoft, on what gets lifted

"Concise answers: One- to two-sentence responses that directly address a question. Structured formatting: Lists, tables, and Q&A blocks that can be lifted cleanly. Strong headings: Signals that help AI know where a complete idea starts and ends. Self-contained phrasing: Sentences that make sense even when pulled out of context."

That last line is the only definition of a quotable sentence any engine has published. Anywhere. In the three years since these products shipped. Has anyone on your team read it?

The same guide describes the mechanism, in a vendor's own words. "AI assistants don't read a page top to bottom like a person would. They break content into smaller, usable pieces," in what Microsoft calls parsing. "These modular pieces are what get ranked and assembled into answers."

Chapter 4 argued that from the outside. Here it is from the inside, by a vendor.

Same picture. Different source.

The guide also contains the only punctuation advice any engine has ever published. It runs to three lines.

Keep punctuation simple. Avoid decorative arrows and symbols, which "break parsing." Be cautious with em dashes, because "overuse can confuse sentence structure for machines."

This book bans em dashes for reasons that have nothing to do with parsers. Take the coincidence for what it is worth, which is not much. No engine has said more.

Bing's Webmaster Guidelines say it again. Formal documentation this time, under a heading that reads "Ensure Content can be Verified Independently."

Their sentence: "URLs are more likely to be selected for grounding queries and citations when content stands on its own."

Then three conditions. "Facts and definitions are explicit." "Key statements do not rely on implied content." "Important information is visible on the URL itself."

That is the whole of Chapter 4, in eight words, written by an engine. Read the middle one twice.

The rest of this chapter leans on that passage. So here is the honesty.

That October guide is a Microsoft Advertising blog post. Not webmaster documentation. No last-updated date.

The Bing Webmaster Tools blog links to it by name as the deeper guidance, and its author co-wrote the AI Performance launch post. So it is adopted rather than official.

Take it as the closest thing to a vendor style guide in existence. Not as a spec.

One more thing before you build on it. Section 10 of Bing's guidelines tells you to "use the data-snippet attribute to specify allowed content." Try it.

No such attribute exists.

Not in Bing's own robots-directives documentation. Not in their data-nosnippet announcement. Not anywhere Microsoft publishes.

So the one published claim that an author can designate which text an engine may quote is an error in a live vendor guideline. A live one.

Somebody should tell them.

The engine that told you not to bother

Google's optimization guide has a mythbusting section. One entry is titled "Rewriting content just for AI systems."

Note the flatness of it: "You don't need to write in a specific way just for generative AI search." No hedge.

The reason follows. AI systems "can understand synonyms and general meanings of what someone is seeking, in order to connect them with content that might not use the same precise words."

The rest of that page, and the helpful-content page it points at, read like an itemized refund. Read them with your last invoice open.

No word count. Their words: "Are you writing to a particular word count because you've heard or read that Google has a preferred word count? (No, we don't.)"

No special files. "You don't need to create new machine readable files, AI text files, markup, or Markdown." No schema. "Structured data isn't required for generative AI search."

Two engines. Microsoft in October 2025, Google seven months later. Opposite instructions.

Before you pick one, notice they answer different questions. Google is denying an obligation. Microsoft is describing what its parser does.

Both can be true at once, and probably are. Nobody is required to write differently. Some sentences still travel better than others.

There is a second reason to hold both. Google does not need your help. It owns the index, the ranking stack and twenty years of query logs.

Microsoft is building a grounding business. So it tells suppliers how to package the goods.

Read each statement as a market position, not a lab result. Not one of them is a lab result.

What neither of them has done is test it. When did anyone last ask them to?

What nobody has published

No engine has ever released a controlled result showing that a sentence rewrite changed whether it was extracted or cited. Nobody has tried in public.

Not one. Not in any direction.

Would you buy a lever nobody has measured? Every claim in the vendor record is asserted rather than measured. Microsoft's included. The gaps are specific enough to list, and the list is the useful part:

  • The extraction unit on any consumer surface. Anthropic publishes sentence-level chunking for its Citations API, over documents you upload. Perplexity publishes span labeling for its Search API. Microsoft publishes passage-level evidence objects for Web IQ, an infrastructure product. Nobody states the unit inside AI Overviews, AI Mode, ChatGPT search or Copilot answers.
  • Whether liftable phrasing earns attribution. Microsoft says assistants "can often lift these pairs word for word." No engine says whether being liftable makes you more likely to be named.
  • Anything about person, tense or voice. Nobody has said whether "researchers found X" beats "we found X." On hedging there is one line, from Microsoft, telling you to "anchor claims in measurable facts." Nobody says what happens if you do not.
  • Reading level. Absent in every direction. No target, no range, and no statement that it does not matter.

Then the one that should bother you most:

Perplexity publishes that spans get labeled vital, irrelevant, duplicate or other. It has never published what makes one vital.

That is the most valuable unwritten document in this field.

So ask your agency one question: what are they using instead?

Four experiments, four different questions

The literature ran the tests the vendors did not. It reaches four different answers. That sounds like a mess. It is not.

Nobody is wrong. They measure four different things.

Table 5.1 · The same tactics, four outcome variablesWhy the results disagree
StudyWhat it measuredRetrieval included?Result on wording
Aggarwal et al.
KDD 2024
Share of the answer attributed to your source, position-weightedNo. Fixed five sourcesQuotations best, 19.3 to 27.2
C-SEO Bench
NeurIPS 2025
Citation rank among the candidatesNo. Fixed candidate list3 of 54 cases significant
Vishwakarma et al.
SIGIR 2026
Which of two sources is cited firstNo. Sources injectedLarge, consistent direction
SAGEO Arena
KDD 2026
Whether you make the candidate list at allYes. Full pipelineEvery body rewrite at or below baseline
Three of the four hand the model a candidate set that already contains you, then ask how prominently you appear in the answer. Only the fourth asks whether you get into the set. That single column explains most of the contradictions in this literature, and it is the column nobody prints when they quote a number at you. Ask for it by name.

So a tactic can win one of these and lose another. One system does exactly that. Yours may be doing it now.

AutoGEO, from Carnegie Mellon and Vody, stopped guessing at what engines reward. It mined the preferences instead.

The result beat the strongest hand-written baseline by 51%. Averaged across three datasets, on the share-of-answer metric. A real result.

Then SAGEO Arena ran a version of the same strategy through a full retrieval pipeline. It came last of ten. Average rank drop of 22 positions.

Best in class at being used. Worst in class at being found.

Is that on your dashboard?

Two caveats, and both are ours: this is the chapter's most quotable comparison, and it is not a knockout.

SAGEO built its own reimplementation rather than running the released system. And AutoGEO was never designed for retrieval survival. It is losing a game it did not enter.

The comparison still stands. Optimize hard for the half you can measure. Nothing warns you about the half you cannot.

Reading level: watch it happen to one variable

Take reading level, the single most confidently prescribed number in this field. Four studies have looked at it.

Aggarwal is quoted at a 15 to 30% lift. Read the sentence. It bundles simplification with fluency, and names no metric.

Simplification alone is 22.0 against a 19.3 baseline. Share of answer, inside a fixed context. So: a lift, on one axis.

C-SEO Bench found Simple Language null in all six domains under GPT-4o-mini, on citation rank, inside a fixed context. Under Claude Haiku it went negative.

SAGEO Arena found it cost 4.18 rank positions, on retrieval, in a full pipeline. A loss, on the other.

And Wan, Wallace and Klein at ACL 2024 found no meaningful correlation between Flesch-Kincaid and how convincing a model found a passage at all. No signal.

Four studies. Four outcome variables. Four answers. Note that this four is not the four in Table 5.1: it swaps Vishwakarma for Wan.

Keep their broader finding, because it applies to everything here. Stylistic features "play a considerably less impactful role in determining the convincingness of text than measures of relevance."

Models, in their words, "largely ignor[e] stylistic features that humans find important."

So when somebody hands you a target reading level, ask which of those four numbers they are quoting. They will not know.

The one result three teams reached separately

Strip out the disagreements and one finding is left standing. Three independent groups, three corpora, no coordination. Do not swap common words for rare or technical ones.

SAGEO Arena, accepted at KDD 2026, pushed ten rewriting strategies through retrieval, reranking and generation. All ten.

Their sentence: "optimizing body text alone consistently degrades visibility across all stages."

Then the specific version. At retrieval, the strategies "that replace common expressions with domain-specific terms (e.g., technical words) or uncommon vocabulary (e.g., unique words) show the largest retrieval drops."

And the mechanism, which is the part worth memorizing. Their words: "replacing terms like 'eating' with 'alimentary routines' or 'sleeping' with 'somnolence' directly reduces term overlap, causing BM25-based retrievers to assign lower relevance scores."

E-GEO, on a 2,000-query test set drawn from 13,747 e-commerce queries, put Technical Terms and Unique Words in its negative bucket. C-SEO Bench found both null in every domain. Negative under Claude Haiku.

Three teams. Three corpora. Same direction. Nobody was looking for it.

Now discount it honestly. The headline number does not survive contact with production.

SAGEO's collapse is a BM25 result. BM25 matches words. Swap the words and you lose, by construction.

Price the finding the way this chapter is about to ask you to price everybody else's. No significance testing appears anywhere in that paper. Single run, no variance reported, a Qwen reranker rather than a commercial engine.

Their own Table 4 reruns it with a dense retriever. The average rank drop falls from 4.54 positions to 0.95. With hybrid retrieval, 2.81.

Real engines run dense or hybrid. What served your last query? Nobody will tell you.

So the cost is real. No retriever average comes back positive. And the size depends on a retriever nobody will tell you about.

Figure 5.2 · What the vocabulary swap costsSame rewrites, three retrievers
AVERAGE RANK CHANGE, BODY-TEXT REWRITES · NEGATIVE IS WORSE 0 Sparse BM25 -4.54 Hybrid -2.81 Dense embeddings -0.95 Word matching. Change the word, lose the match. Both signals. Two thirds of the penalty. Meaning matching. The penalty shrinks, and stays a penalty.
The same ten body-text rewrites, run through three retrieval architectures inside one experiment. Under sparse retrieval the vocabulary swap is expensive, because sparse retrieval is word overlap and nothing else. Under dense retrieval it costs about a fifth as much. What never happens, at any retriever, is a gain on average. Nobody publishes which retriever ran your query, so plan on the small number and act on the direction.

The table the field is about to quote wrong

One 2026 study went at wording head-on. It is the best-designed thing in this chapter's evidence base. Vishwakarma and colleagues at Sprinklr, published at SIGIR 2026, ran 252,000 trials.

Eighteen content factors. Six models. Two candidate documents at a time, order counterbalanced, brand names stripped out. The outcome is narrow and clean: which of the two gets the first citation.

Their hedging manipulation, verbatim. Look at the swing they tested.

Confident: "The CleanBot Aroma Pro X3 delivers exceptional cleaning with 30,000 Pa suction and 99.99% germ elimination." That is variant A.

Hedged: "The CleanBot Aroma Pro X3 might possibly deliver cleaning with what could be around 30,000 Pa suction." Nobody writes the second version on purpose. Plenty of legal reviews produce it anyway.

Four factors cleared every model with very large effects. The authors call them gatekeepers: "Topic Mismatch, Price Not Mentioned, Recent vs Old Timestamp, and Lower List Position." Four. That is the whole list.

Specifications present, evidence attached and query terms present all ran the same direction in all six. Confident beat hedged in every one.

A second team found it on a different task, without looking for it.

Van de Sande and colleagues at Radboud took verified-false claims and rewrote them with uncertainty markers. Meaning held fixed, every rewrite checked by hand. Then they asked three models to fact-check them.

Hedging flipped the classification from false to not-false in 25% of cases. Framing the same claim as a belief, "I believe X," flipped it in 50 to 56%. Half the time.

That is a fact-checking task, not a citation task. But it says something the Sprinklr study cannot.

Hedging does not just make a claim less attractive. It changes what the model thinks the claim is.

Which is the argument to take into your next legal review. A hedge is not a weaker version of the claim. It is a different claim.

And one null worth having: "Formatting choices (Content Structure, Scattered Information) had no impact, suggesting LLMs parse content regardless of visual organization."

The biggest wording number anyone can quote off that table is 754, for confident language under Claude. It will be in a hundred posts by Christmas. The paper does not print it in bold. In their notation, bold means significant.

Their own footnote: "Non-bold ORs lack reliable evidence of an effect."

The other headline figures print as ">10k." Not an effect size either.

The estimates become very large "under quasi-separation." They "treat them as indicating a decisive win for variant A rather than a finely resolved numeric ratio."

A verdict, not a multiplier. No forecast should rest on one.

The asterisks and daggers through that table are not significance stars. The legend says "Degenerate Hessian" and "Singular fit." Fitting failures, both.

A cell carrying one is a cell whose error bars nobody should trust. So the direction is solid across six models. The magnitudes are not numbers.

Two different things, printed in the same table.

Anyone who quotes you a multiple off that table has not read the footnote. How many of the studies on your last strategy deck did you check that far?

Its limits deserve the same honesty. No search engine is ever called. Sources are injected into the context as a fabricated tool response.

Every document is a product review blog. The corpus was generated end to end by another model.

Which is a problem about where one number came from. This field has a bigger one.

Where the prescriptions actually come from

Chapter 4 caught the field quoting a developer API as though it described web retrieval. That was not an isolated incident.

Here it is again, inside the most repeated piece of writing advice in the discipline. Second time in this book.

Table 5.3 · Traced to sourceFive prescriptions, five origins
What you were toldWhere it actually comes fromVerdict
"Answer in the first 40 to 60 words"Two studies of where Google's featured snippet box truncates text, the later one 7,854 keywords, desktop only, in 2021A display constraint, relabeled
"Chunk to a fixed token count"OpenAI's file search defaults, 800 tokens with 400 overlap, for documents you uploadSame error as Chapter 4
"Write at grade 8"Plain-language convention, imported wholeNo AI search study produces the number
"2.8x more citations from heading hierarchy"One vendor study, 12,000 URLs. 68.7% of ChatGPT-cited pages used sequential headings against 23.9% of Google's top resultsA prevalence ratio, not a lift
"Add llms.txt"A 2024 proposal for assembling context for coding assistantsRefuted, see below
Every row is a real artifact repurposed as a ranking factor. The first is the instructive case: 40 to 60 words measures a 2021 search results page, not a property of language models, and the same study produced the "keep paragraphs to two or three sentences" rule. Two of the most repeated instructions in AI content writing are one line of research about the size of a box on a screen.

The llms.txt row deserves its own paragraph. Not for the reason you think. This is not an absence of evidence.

It is evidence of absence, from four independent designs.

Who is still billing you for the file?

Jeremy Howard proposed the file in September 2024. The purpose was helping developers assemble context for coding assistants reading library documentation.

Search visibility was never the point. What happened next, in four separate designs.

A study across roughly 300,000 domains found no relationship. Removing the feature improved their model's accuracy.

A server-log analysis across 137,000 domains found 97% of these files got zero requests in a month.

A before-and-after across ten sites found no measurable change on eight. And a single-site log study found the file drew 84 visits, against about 265 for an average content page.

Google has said twice that it does not use it.

John Mueller put it best: "you can tell when you look at your server logs that they don't even check for it."

Check yours before you argue with him.

Nobody else has said they do. Whose backlog is it still on?

One more, because it shows the machinery rather than a single bad number.

The largest dataset on question-form headings covers 216,524 pages. Pages using them averaged 3.4 citations, against 4.3 for straightforward headings.

Worse, not better. Which is how that number gets quoted.

Now read the very next paragraph. Their model treats the absence of an FAQ section as a negative signal.

And their own recommendation list tells you to use question-based titles, reporting almost seven times the impact on smaller domains.

One study. Two results. Opposite directions.

Both halves are in circulation, each quoted by people who never print the other. And this book has to live by the rule it just applied to Sprinklr.

So price them both. What would either one have to show to change your template?

3.4 against 4.3 is an uncontrolled mean with no dispersion published.

The model result is a feature attribution, not an experiment.

Neither is evidence you should restructure a site on. That is the finding, and it is duller than either headline. Which is usually the tell.

The tone question, since somebody is billing you for it

Authoritative rewriting is the oldest tactic in the category, and the paper that invented the category tested it first: with the result nobody quotes.

Aggarwal's own words: "we find no significant improvement, demonstrating that Generative Engines are already somewhat robust to such changes."

C-SEO Bench found it null in every domain under GPT-4o-mini, and negative under Claude Haiku. E-GEO, in its current version, has it negative across every re-ranker it tested.

It is not nothing. On share of answer, inside a fixed context, it scores 21.3 against a 19.3 baseline. AutoGEO's numbers run the same way.

But every one of those positives is measured after retrieval. Nobody has shown tone getting you into the pool.

And the one study that looked at retrieval put every body rewrite at or below baseline.

So the honest instruction is narrower than the pitch: do not buy tone as a visibility lever. Buy it, if you buy it, because you want to sound like that. Not because it moves anything.

The method Chapter 1 promised you

Chapter 1 gave you four questions to ask of a passage. Chapter 4 gave you the joins where a passage breaks.

This is where it becomes a rewrite you can hand to somebody.

The whole method is one distinction. Everything above is why it holds.

Before it, one disclosure. Moves one to three are Chapter 1 and Chapter 4 restated in a form you can hand over. Move four is the dating rule Chapter 3 promised you and left here. Move six is Chapter 1's pricing argument with a number behind it.

Only move five is new.

The check is not new either. Chapter 2 got there first. It argued that your buyer's vocabulary beats your own.

What is new is the price tag. Chapter 2 framed buyer vocabulary as something to add. This chapter prices what it costs when somebody removes it. On your invoice. In a rewrite you approved.

A sentence has two layers, and only one of them is safe to edit. One layer per gate, in the order Figure 5.0 puts them.

The claim layer is what your sentence asserts. How firmly, about whom, on what evidence, as of when.

Almost every controlled result in this chapter that came back positive lives there.

The vocabulary layer is the words you chose to say it in.

Almost every controlled result that came back negative lives there.

So edit the claim. Leave your words alone.

Which is the reverse of what almost every content service sells. So ask yours: which layer do you edit?

Rewriting vocabulary is the part that is easy to sell. It is also the part that shows up in a before-and-after screenshot. You have seen that deck.

Artifact 5.4 · The two-layer rewriteSix moves, then one check
  1. Name the subject inside the sentence. Not in the heading above it. Not in the paragraph before it. If the sentence opens with "It," "This" or "Our platform," the subject is sitting somewhere the reader may never receive.
  2. Delete the unbounded qualifier, or replace it with a number. "Usually," "typically," "most," "fast," "industry-leading." An unbounded qualifier is not a cautious claim. It is no claim, and a model that cannot state it cannot ground on it.
  3. Attach the condition to the number. A figure without its unit, its population and its date is a figure somebody else has to caveat for you. "6 working days" is weaker than "6 working days, median, for a 50-seat rollout."
  4. Date anything that can go stale, in the prose. Recency was one of only four factors clearing every model in the SIGIR study. Put it in the sentence, not only in the byline, because the byline may not travel with the sentence.
  5. Say who verified it, if anyone did. Evidence attached ran ahead of evidence absent in all six models, significantly in four. An audit, a test, a certification, a sample size. One clause is enough.
  6. Put a number on the price, or a range. "Price Not Mentioned" was another of those four gatekeepers, and it is the one most B2B pages fail. Chapter 1 already answered your legal team: bound the range, date the claim, attribute it to a named source.
Then the check, and this is the one people skip. Take the first sentence under each H2 of your twenty highest-revenue pages. For each one, list every content noun and verb the rewrite introduced, excluding anything inside a number's condition clause, because those words carry the claim rather than the match. Then export twelve months of Search Console queries, or pull your internal site-search log, and look each word up. A word that appears zero times is new to your buyers, so put the original back. Three or more new content words in one sentence is a rewrite to reject outright.
Then the number. Score each audited sentence 1 if it passes the check and 0 if it fails, and report the percentage. Under 80 and your brief is buying tone. Eighty is ours: no engine publishes a threshold, and this one sits where a sample stops looking like a few bad sentences and starts looking like a house instruction. Re-run on twenty fresh sentences after the brief changes, same scorer, because nothing else here gives you a second reading.
What this does not measure. It counts how much of your rewriting budget is being spent against gate one. It predicts nothing about citations, and it cannot tell you what any real segmenter did with any of it. Chapter 2's cosine test answers the neighboring question, whether a rewrite is a genuine reframing or a synonym. Run that after this one, not instead of it.
Six moves that touch what the sentence claims, and one check that protects the words it claims it in. Moves one to three are Chapters 1 and 4 in a form you can delegate. The check is the part with a data source, a threshold and a re-run behind it, which is deliberate: it is the one step where two reviewers would otherwise disagree. What it hands back is a single percentage, and that percentage is the only number in this chapter you generated yourself.

Here is the same edit run on a whole sentence. And on the version an agency would ship instead.

Figure 5.5 · One claim, two layersIllustrative · the two-layer edit, and the one to reject
BEFORE
Our vertically integrated fulfillment architecture facilitates expedited order disposition, typically within a compressed timeframe.
Fourteen words. No subject, no number, no date. And every content word is one a buyer would never type. Nobody searches for "order disposition." They search for shipping time.
THE REWRITE YOU WILL BE SOLD
Acme Logistics operates an integrated fulfillment capability that dispatches 94% of available inventory within the same business day, subject to a 2pm local cut-off, with non-stocked SKUs following in 3 to 5 working days. Both figures are 2025 medians across 41,000 orders.
Claim layer fixed. Subject named, numbers bounded, conditions attached. Vocabulary layer elevated: dispatches, available inventory, business day, non-stocked SKUs. Gate two opened. Gate one narrowed. Nothing in your reporting separates this from the version below.
AFTER
Acme Logistics ships 94% of in-stock orders the same day when the order lands before 2pm local time. Out-of-stock items ship in 3 to 5 working days. Both figures are 2025 medians across 41,000 orders.
Subject named. Two bounded claims. Every number carries its condition, and both are dated. And the words are the words a buyer types: ships, in stock, same day, working days.
The middle panel is the one to study. Its claim layer is as good as the bottom panel's, and every reporting line you have would score them the same. It also swapped four phrases a buyer types for four a buyer never would, which is the only edit here that costs anything at gate one. Notice too what the bottom version costs: duller read straight through, and two numbers committed in public. That is the whole price of being quotable, and no version avoids it.

The objection this chapter has to answer

"You told me not to upgrade my vocabulary, because sparse retrieval loses term overlap. Then you told me real engines run dense, where the cost is 0.95 rank positions. From a single-run study with no significance testing.

"Your headline finding is a rounding error on the retriever that actually serves my query."

That is the right objection. Sharper than the one about the vendors disagreeing.

Three answers. The direction is unanimous across three teams, and two of them never used BM25 at all. The 0.95 is an average, and the tactics being sold to you sit at the bad end of the spread. Not the middle.

And the size is not really the point. The point is that the trade is invisible. Whatever it costs, you are paying for the side of it that loses. No report you own separates the two.

Then the wider objection, because it is coming anyway. Two engines disagree. Neither ran an experiment. Four studies measure four things. The best of them prints numbers it tells you not to trust.

Why act on any of it? Table 5.3 shows you what the alternative is. Worse.

Take that objection at full strength too. This is the thinnest evidence base in the book, and it is not close. By some way.

Chapter 3 rests on vendor documentation you can verify with curl in an afternoon.

This chapter rests on four studies that disagree, and one style guide from an advertising blog.

So the method above is built to be cheap and to fail safe.

Naming your subject, bounding your numbers and dating your claims cost a writer twenty minutes. They improve the page for humans whatever the engines turn out to do. That is the hedge you want.

And the check costs nothing at all: it is an instruction to stop doing something. How often do you get one of those?

Nothing in Artifact 5.4 asks you to believe a single number in this chapter. That is deliberate. It is also how to treat any advice about a system nobody has documented.

What is left: six instructions

The engines have published almost nothing about writing. What little they have published contradicts itself. The studies measure prominence inside a context you already reached. All but one. That one measures whether you reach it. It says every rewrite costs you something.

Against all that, six things hold.

Name your subject. Bound your claims. Date them. Say who checked them. Put a number on the price. Do not upgrade your vocabulary.

Six instructions, and your writers can hold all six in their heads: subject, bound, date, verify, price, and leave the words alone.

That is the only reason any of them will survive a quarter.

The first five are unglamorous and free. The last one requires canceling something.

Which is why it is the one worth an argument in your next content review.

One loose end first, and Chapter 11 picks it up. Nothing in your reporting separates a rewrite that helped from one that quietly cost you.

That chapter exists for exactly this.

Now notice what the SIGIR study put at the top of its list. Above every wording factor it tested.

Topic mismatch. Not phrasing. Whether your page is about the thing at all.

Which is a question about what your page is about, and that is not a wording problem. It is a bigger one.

Chapter 6 is where it gets solved.

Chapter 05 · What to do with this:
  • Edit the claim, not the words. Then run the check: every content noun and verb you introduced, looked up in your query data.
  • Date anything that can go stale, in the prose. Recency cleared all six models, and the byline may not travel with the sentence.
  • Put a number on the price. A missing price cleared every model as a gatekeeper, and most B2B pages fail it.
  • Ask any number for its outcome variable. Prominence inside a context you reached is not the same result as reaching it.
  • Do not buy tone as a lever. Null in the paper that invented it, negative in two more, every positive after retrieval.
  • Cancel the llms.txt ticket. Four studies, no effect. 97% of the files get zero requests, and Google has twice said no.
Sources
  • Krishna Madhavan, Microsoft Advertising, "Optimizing Your Content for Inclusion in AI Search Answers," 8 October 2025. The four snippet-eligibility criteria, quoted in part. A Microsoft Advertising blog post, linked by the Bing Webmaster Tools blog as the deeper guidance, carrying no last-updated date.
  • Bing Webmaster Guidelines, sections 10 and 15 to 18. Grounding and citation guidance, entity naming, single topic per URL, key information early. The section 10 instruction to use a "data-snippet" attribute has no corresponding documentation anywhere Microsoft publishes.
  • Google Search Central, "Optimizing your website for generative AI features on Google Search," published 15 May 2026, updated 10 July 2026. Mythbusting entries on rewriting for AI, machine-readable files and structured data.
  • Google Search Central, "Creating helpful, reliable, people-first content," updated 10 December 2025. The word-count question, verbatim.
  • Perplexity, "Architecting and Evaluating an AI-First Search API," 25 September 2025, and "Search API: better extraction, dynamic benchmarks," 11 March 2026. Span labeling published, criteria never published.
  • Anthropic, Claude Platform citations documentation. Sentence-level chunking, for developer-supplied documents only. OpenAI publishes nothing to website owners about text.
  • Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, "GEO: Generative Engine Optimization," KDD 2024, arXiv:2311.09735. Position-adjusted word count, baseline 19.3, Quotation Addition 27.2, Easy-to-Understand 22.0, Authoritative 21.3. Sources fixed at the top five results, so retrieval is held constant. Authoritative rewriting scored 21.3 against the 19.3 baseline, and the paper's own gloss is that "we find no significant improvement, demonstrating that Generative Engines are already somewhat robust to such changes."
  • Puerto, Gubri, Green, Oh and Yun, "C-SEO Bench: Does Conversational SEO Work?", NeurIPS 2025 Datasets and Benchmarks Track. Three of 54 cases significant. One of the two effective methods is a markdown summary rather than a rewrite, and the other is defined as all eight style transformations applied at once plus structural formatting.
  • Kim, Jeong, Kim, Lee and Lee, "SAGEO Arena," KDD 2026, arXiv:2602.12187. Ten rewriting strategies. Body-text-only average rank change of 4.54 positions under BM25, 2.81 hybrid, 0.95 dense, from their Table 4. The AutoGEO-style strategy placed tenth of ten at 22.35 positions. No significance testing, no variance and no repeated runs are reported anywhere in the paper, and the pipeline is BM25 with a Qwen reranker rather than a commercial engine.
  • Wu, Zhong, Kim and Xiong, "What Generative Search Engines Like and How to Optimize Web Content Cooperatively" (AutoGEO), ICLR 2026, arXiv:2510.11438. A 50.99% average gain over Fluency Optimization across three datasets, on the share-of-answer metric, over a fixed five-document candidate set. The version SAGEO Arena evaluated is a prompt reimplementation rather than the released system, which is why this chapter reports the comparison and not a verdict.
  • Bagga, Farias, Korkotashvili, Peng and Wu, "E-GEO: A Testbed for Generative Engine Optimization in E-Commerce," arXiv:2511.20867, version 2, 14 July 2026. 13,747 queries. Technical and Unique in the negative bucket of fifteen hand-written heuristics, and Authoritative negative across all five re-rankers. Unique turns positive after the paper's own prompt optimization. Technical only reaches about zero, and stays negative on three of the five re-rankers, which is a result about prompts rather than about wording.
  • Wan, Wallace and Klein, "What Evidence Do Language Models Find Convincing?", ACL 2024. ConflictingQA, 238 questions, 2,208 retrieved paragraphs, five models. Flesch-Kincaid and unique-token count show no correlation with judged convincingness. Reported as figures rather than coefficients, so quote the sentences and not a number.
  • Van de Sande, Acar, van Woudenberg and Larson, Radboud University, "On Fact and Frequency," arXiv:2503.04271. Verified-false claims rewritten with uncertainty markers, meaning held fixed and every rewrite hand-checked, then fact-checked by three models. Hedging flipped 25% of classifications from false to not-false. Belief framing flipped 50 to 56%. A fact-checking task, not a citation task.
  • AirOps, "Structuring content for LLMs." 12,000 URLs across 900 queries and 15 industries. 68.7% of ChatGPT-cited pages used sequential heading structure against 23.9% of Google's top results, which is a prevalence ratio between two corpora and not a measured citation lift.
  • Microsoft, "Announcing Microsoft Web IQ," 2 June 2026. Passages and structured evidence objects, in a grounding infrastructure product rather than a consumer surface.
  • Vishwakarma, Kumar and Jamidar, "What Gets Cited: Competitive GEO in AI Answer Engines," SIGIR 2026, arXiv:2605.25517. 252,000 trials, 18 factors, six models, two documents per trial. Outcome is which source receives the first citation. Bold denotes significance, ">10k" denotes quasi-separation rather than an effect size, and the asterisk and dagger denote degenerate Hessian and singular fit. No multiple-comparison correction is reported. The corpus is product review blogs, generated by another model.
  • Evan Hall, Portent, "Featured snippet display lengths," 3 June 2021. 7,854 featured-snippet keywords, desktop only. It replicates an earlier Ghergich and SEMrush finding on word count, and reports the two to three sentence result separately. Between them they are the origin of both circulating rules.
  • Jeremy Howard, Answer.AI, "The /llms.txt file," 3 September 2024. John Mueller on Reddit, 17 April 2025, and Gary Illyes at Search Central Deep Dive APAC, 23 July 2025. SE Ranking, roughly 300,000 domains, 7 November 2025. Server-log analysis across 137,000 domains, May 2026. Otterly.ai, 5 February 2026. Previsible, ten sites, 20 January 2026.
  • SE Ranking, 129,000 domains and 216,524 pages across 20 niches, 24 November 2025. Question-style headings averaged 3.4 citations against 4.3 for straightforward headings. The same article reports a SHAP attribution on FAQ presence running the other way, and recommends question-based titles, reporting close to seven times the effect on smaller domains. Both figures are descriptive, and neither is an experiment.
Now taking new clients · limited spots

Reading it is one thing. Running it is another.

We build and operate this system for a small number of clients each quarter. Book a session and we'll audit where you currently sit in the pipeline.

See case studies Book Strategy Session →