Two search engines have published writing guidance for AI answers. They contradict each other. Neither has run a single experiment.
The style guide nobody quotes
Exactly one engine has ever told you how to write a sentence for extraction. It was not Google.
Microsoft published it on 8 October 2025, under the title "Optimizing Your Content for Inclusion in AI Search Answers." Four criteria, under a heading asking what makes content eligible for featured snippets. Four.
Nobody quotes them. Not one agency deck.
"Concise answers: One- to two-sentence responses that directly address a question. Structured formatting: Lists, tables, and Q&A blocks that can be lifted cleanly. Strong headings: Signals that help AI know where a complete idea starts and ends. Self-contained phrasing: Sentences that make sense even when pulled out of context."
That last line is the only definition of a quotable sentence any engine has published. Anywhere. In the three years since these products shipped. Has anyone on your team read it?
The same guide describes the mechanism, in a vendor's own words. "AI assistants don't read a page top to bottom like a person would. They break content into smaller, usable pieces," in what Microsoft calls parsing. "These modular pieces are what get ranked and assembled into answers."
Chapter 4 argued that from the outside. Here it is from the inside, by a vendor.
Same picture. Different source.
The guide also contains the only punctuation advice any engine has ever published. It runs to three lines.
Keep punctuation simple. Avoid decorative arrows and symbols, which "break parsing." Be cautious with em dashes, because "overuse can confuse sentence structure for machines."
This book bans em dashes for reasons that have nothing to do with parsers. Take the coincidence for what it is worth, which is not much. No engine has said more.
Bing's Webmaster Guidelines say it again. Formal documentation this time, under a heading that reads "Ensure Content can be Verified Independently."
Their sentence: "URLs are more likely to be selected for grounding queries and citations when content stands on its own."
Then three conditions. "Facts and definitions are explicit." "Key statements do not rely on implied content." "Important information is visible on the URL itself."
That is the whole of Chapter 4, in eight words, written by an engine. Read the middle one twice.
The rest of this chapter leans on that passage. So here is the honesty.
That October guide is a Microsoft Advertising blog post. Not webmaster documentation. No last-updated date.
The Bing Webmaster Tools blog links to it by name as the deeper guidance, and its author co-wrote the AI Performance launch post. So it is adopted rather than official.
Take it as the closest thing to a vendor style guide in existence. Not as a spec.
One more thing before you build on it. Section 10 of Bing's guidelines tells you to "use the data-snippet attribute to specify allowed content." Try it.
No such attribute exists.
Not in Bing's own robots-directives documentation. Not in their data-nosnippet announcement. Not anywhere Microsoft publishes.
So the one published claim that an author can designate which text an engine may quote is an error in a live vendor guideline. A live one.
Somebody should tell them.
The engine that told you not to bother
Google's optimization guide has a mythbusting section. One entry is titled "Rewriting content just for AI systems."
Note the flatness of it: "You don't need to write in a specific way just for generative AI search." No hedge.
The reason follows. AI systems "can understand synonyms and general meanings of what someone is seeking, in order to connect them with content that might not use the same precise words."
The rest of that page, and the helpful-content page it points at, read like an itemized refund. Read them with your last invoice open.
No word count. Their words: "Are you writing to a particular word count because you've heard or read that Google has a preferred word count? (No, we don't.)"
No special files. "You don't need to create new machine readable files, AI text files, markup, or Markdown." No schema. "Structured data isn't required for generative AI search."
Two engines. Microsoft in October 2025, Google seven months later. Opposite instructions.
Before you pick one, notice they answer different questions. Google is denying an obligation. Microsoft is describing what its parser does.
Both can be true at once, and probably are. Nobody is required to write differently. Some sentences still travel better than others.
There is a second reason to hold both. Google does not need your help. It owns the index, the ranking stack and twenty years of query logs.
Microsoft is building a grounding business. So it tells suppliers how to package the goods.
Read each statement as a market position, not a lab result. Not one of them is a lab result.
What neither of them has done is test it. When did anyone last ask them to?
What nobody has published
No engine has ever released a controlled result showing that a sentence rewrite changed whether it was extracted or cited. Nobody has tried in public.
Not one. Not in any direction.
Would you buy a lever nobody has measured? Every claim in the vendor record is asserted rather than measured. Microsoft's included. The gaps are specific enough to list, and the list is the useful part:
- The extraction unit on any consumer surface. Anthropic publishes sentence-level chunking for its Citations API, over documents you upload. Perplexity publishes span labeling for its Search API. Microsoft publishes passage-level evidence objects for Web IQ, an infrastructure product. Nobody states the unit inside AI Overviews, AI Mode, ChatGPT search or Copilot answers.
- Whether liftable phrasing earns attribution. Microsoft says assistants "can often lift these pairs word for word." No engine says whether being liftable makes you more likely to be named.
- Anything about person, tense or voice. Nobody has said whether "researchers found X" beats "we found X." On hedging there is one line, from Microsoft, telling you to "anchor claims in measurable facts." Nobody says what happens if you do not.
- Reading level. Absent in every direction. No target, no range, and no statement that it does not matter.
Then the one that should bother you most:
Perplexity publishes that spans get labeled vital, irrelevant, duplicate or other. It has never published what makes one vital.
That is the most valuable unwritten document in this field.
So ask your agency one question: what are they using instead?
Four experiments, four different questions
The literature ran the tests the vendors did not. It reaches four different answers. That sounds like a mess. It is not.
Nobody is wrong. They measure four different things.
| Study | What it measured | Retrieval included? | Result on wording |
|---|---|---|---|
| Aggarwal et al. KDD 2024 | Share of the answer attributed to your source, position-weighted | No. Fixed five sources | Quotations best, 19.3 to 27.2 |
| C-SEO Bench NeurIPS 2025 | Citation rank among the candidates | No. Fixed candidate list | 3 of 54 cases significant |
| Vishwakarma et al. SIGIR 2026 | Which of two sources is cited first | No. Sources injected | Large, consistent direction |
| SAGEO Arena KDD 2026 | Whether you make the candidate list at all | Yes. Full pipeline | Every body rewrite at or below baseline |
So a tactic can win one of these and lose another. One system does exactly that. Yours may be doing it now.
AutoGEO, from Carnegie Mellon and Vody, stopped guessing at what engines reward. It mined the preferences instead.
The result beat the strongest hand-written baseline by 51%. Averaged across three datasets, on the share-of-answer metric. A real result.
Then SAGEO Arena ran a version of the same strategy through a full retrieval pipeline. It came last of ten. Average rank drop of 22 positions.
Best in class at being used. Worst in class at being found.
Is that on your dashboard?
Two caveats, and both are ours: this is the chapter's most quotable comparison, and it is not a knockout.
SAGEO built its own reimplementation rather than running the released system. And AutoGEO was never designed for retrieval survival. It is losing a game it did not enter.
The comparison still stands. Optimize hard for the half you can measure. Nothing warns you about the half you cannot.
Reading level: watch it happen to one variable
Take reading level, the single most confidently prescribed number in this field. Four studies have looked at it.
Aggarwal is quoted at a 15 to 30% lift. Read the sentence. It bundles simplification with fluency, and names no metric.
Simplification alone is 22.0 against a 19.3 baseline. Share of answer, inside a fixed context. So: a lift, on one axis.
C-SEO Bench found Simple Language null in all six domains under GPT-4o-mini, on citation rank, inside a fixed context. Under Claude Haiku it went negative.
SAGEO Arena found it cost 4.18 rank positions, on retrieval, in a full pipeline. A loss, on the other.
And Wan, Wallace and Klein at ACL 2024 found no meaningful correlation between Flesch-Kincaid and how convincing a model found a passage at all. No signal.
Four studies. Four outcome variables. Four answers. Note that this four is not the four in Table 5.1: it swaps Vishwakarma for Wan.
Keep their broader finding, because it applies to everything here. Stylistic features "play a considerably less impactful role in determining the convincingness of text than measures of relevance."
Models, in their words, "largely ignor[e] stylistic features that humans find important."
So when somebody hands you a target reading level, ask which of those four numbers they are quoting. They will not know.
The one result three teams reached separately
Strip out the disagreements and one finding is left standing. Three independent groups, three corpora, no coordination. Do not swap common words for rare or technical ones.
SAGEO Arena, accepted at KDD 2026, pushed ten rewriting strategies through retrieval, reranking and generation. All ten.
Their sentence: "optimizing body text alone consistently degrades visibility across all stages."
Then the specific version. At retrieval, the strategies "that replace common expressions with domain-specific terms (e.g., technical words) or uncommon vocabulary (e.g., unique words) show the largest retrieval drops."
And the mechanism, which is the part worth memorizing. Their words: "replacing terms like 'eating' with 'alimentary routines' or 'sleeping' with 'somnolence' directly reduces term overlap, causing BM25-based retrievers to assign lower relevance scores."
E-GEO, on a 2,000-query test set drawn from 13,747 e-commerce queries, put Technical Terms and Unique Words in its negative bucket. C-SEO Bench found both null in every domain. Negative under Claude Haiku.
Three teams. Three corpora. Same direction. Nobody was looking for it.
Now discount it honestly. The headline number does not survive contact with production.
SAGEO's collapse is a BM25 result. BM25 matches words. Swap the words and you lose, by construction.
Price the finding the way this chapter is about to ask you to price everybody else's. No significance testing appears anywhere in that paper. Single run, no variance reported, a Qwen reranker rather than a commercial engine.
Their own Table 4 reruns it with a dense retriever. The average rank drop falls from 4.54 positions to 0.95. With hybrid retrieval, 2.81.
Real engines run dense or hybrid. What served your last query? Nobody will tell you.
So the cost is real. No retriever average comes back positive. And the size depends on a retriever nobody will tell you about.
The table the field is about to quote wrong
One 2026 study went at wording head-on. It is the best-designed thing in this chapter's evidence base. Vishwakarma and colleagues at Sprinklr, published at SIGIR 2026, ran 252,000 trials.
Eighteen content factors. Six models. Two candidate documents at a time, order counterbalanced, brand names stripped out. The outcome is narrow and clean: which of the two gets the first citation.
Their hedging manipulation, verbatim. Look at the swing they tested.
Confident: "The CleanBot Aroma Pro X3 delivers exceptional cleaning with 30,000 Pa suction and 99.99% germ elimination." That is variant A.
Hedged: "The CleanBot Aroma Pro X3 might possibly deliver cleaning with what could be around 30,000 Pa suction." Nobody writes the second version on purpose. Plenty of legal reviews produce it anyway.
Four factors cleared every model with very large effects. The authors call them gatekeepers: "Topic Mismatch, Price Not Mentioned, Recent vs Old Timestamp, and Lower List Position." Four. That is the whole list.
Specifications present, evidence attached and query terms present all ran the same direction in all six. Confident beat hedged in every one.
A second team found it on a different task, without looking for it.
Van de Sande and colleagues at Radboud took verified-false claims and rewrote them with uncertainty markers. Meaning held fixed, every rewrite checked by hand. Then they asked three models to fact-check them.
Hedging flipped the classification from false to not-false in 25% of cases. Framing the same claim as a belief, "I believe X," flipped it in 50 to 56%. Half the time.
That is a fact-checking task, not a citation task. But it says something the Sprinklr study cannot.
Hedging does not just make a claim less attractive. It changes what the model thinks the claim is.
Which is the argument to take into your next legal review. A hedge is not a weaker version of the claim. It is a different claim.
And one null worth having: "Formatting choices (Content Structure, Scattered Information) had no impact, suggesting LLMs parse content regardless of visual organization."
The biggest wording number anyone can quote off that table is 754, for confident language under Claude. It will be in a hundred posts by Christmas. The paper does not print it in bold. In their notation, bold means significant.
Their own footnote: "Non-bold ORs lack reliable evidence of an effect."
The other headline figures print as ">10k." Not an effect size either.
The estimates become very large "under quasi-separation." They "treat them as indicating a decisive win for variant A rather than a finely resolved numeric ratio."
A verdict, not a multiplier. No forecast should rest on one.
The asterisks and daggers through that table are not significance stars. The legend says "Degenerate Hessian" and "Singular fit." Fitting failures, both.
A cell carrying one is a cell whose error bars nobody should trust. So the direction is solid across six models. The magnitudes are not numbers.
Two different things, printed in the same table.
Anyone who quotes you a multiple off that table has not read the footnote. How many of the studies on your last strategy deck did you check that far?
Its limits deserve the same honesty. No search engine is ever called. Sources are injected into the context as a fabricated tool response.
Every document is a product review blog. The corpus was generated end to end by another model.
Which is a problem about where one number came from. This field has a bigger one.
Where the prescriptions actually come from
Chapter 4 caught the field quoting a developer API as though it described web retrieval. That was not an isolated incident.
Here it is again, inside the most repeated piece of writing advice in the discipline. Second time in this book.
| What you were told | Where it actually comes from | Verdict |
|---|---|---|
| "Answer in the first 40 to 60 words" | Two studies of where Google's featured snippet box truncates text, the later one 7,854 keywords, desktop only, in 2021 | A display constraint, relabeled |
| "Chunk to a fixed token count" | OpenAI's file search defaults, 800 tokens with 400 overlap, for documents you upload | Same error as Chapter 4 |
| "Write at grade 8" | Plain-language convention, imported whole | No AI search study produces the number |
| "2.8x more citations from heading hierarchy" | One vendor study, 12,000 URLs. 68.7% of ChatGPT-cited pages used sequential headings against 23.9% of Google's top results | A prevalence ratio, not a lift |
| "Add llms.txt" | A 2024 proposal for assembling context for coding assistants | Refuted, see below |
The llms.txt row deserves its own paragraph. Not for the reason you think. This is not an absence of evidence.
It is evidence of absence, from four independent designs.
Who is still billing you for the file?
Jeremy Howard proposed the file in September 2024. The purpose was helping developers assemble context for coding assistants reading library documentation.
Search visibility was never the point. What happened next, in four separate designs.
A study across roughly 300,000 domains found no relationship. Removing the feature improved their model's accuracy.
A server-log analysis across 137,000 domains found 97% of these files got zero requests in a month.
A before-and-after across ten sites found no measurable change on eight. And a single-site log study found the file drew 84 visits, against about 265 for an average content page.
Google has said twice that it does not use it.
John Mueller put it best: "you can tell when you look at your server logs that they don't even check for it."
Check yours before you argue with him.
Nobody else has said they do. Whose backlog is it still on?
One more, because it shows the machinery rather than a single bad number.
The largest dataset on question-form headings covers 216,524 pages. Pages using them averaged 3.4 citations, against 4.3 for straightforward headings.
Worse, not better. Which is how that number gets quoted.
Now read the very next paragraph. Their model treats the absence of an FAQ section as a negative signal.
And their own recommendation list tells you to use question-based titles, reporting almost seven times the impact on smaller domains.
One study. Two results. Opposite directions.
Both halves are in circulation, each quoted by people who never print the other. And this book has to live by the rule it just applied to Sprinklr.
So price them both. What would either one have to show to change your template?
3.4 against 4.3 is an uncontrolled mean with no dispersion published.
The model result is a feature attribution, not an experiment.
Neither is evidence you should restructure a site on. That is the finding, and it is duller than either headline. Which is usually the tell.
The tone question, since somebody is billing you for it
Authoritative rewriting is the oldest tactic in the category, and the paper that invented the category tested it first: with the result nobody quotes.
Aggarwal's own words: "we find no significant improvement, demonstrating that Generative Engines are already somewhat robust to such changes."
C-SEO Bench found it null in every domain under GPT-4o-mini, and negative under Claude Haiku. E-GEO, in its current version, has it negative across every re-ranker it tested.
It is not nothing. On share of answer, inside a fixed context, it scores 21.3 against a 19.3 baseline. AutoGEO's numbers run the same way.
But every one of those positives is measured after retrieval. Nobody has shown tone getting you into the pool.
And the one study that looked at retrieval put every body rewrite at or below baseline.
So the honest instruction is narrower than the pitch: do not buy tone as a visibility lever. Buy it, if you buy it, because you want to sound like that. Not because it moves anything.
The method Chapter 1 promised you
Chapter 1 gave you four questions to ask of a passage. Chapter 4 gave you the joins where a passage breaks.
This is where it becomes a rewrite you can hand to somebody.
The whole method is one distinction. Everything above is why it holds.
Before it, one disclosure. Moves one to three are Chapter 1 and Chapter 4 restated in a form you can hand over. Move four is the dating rule Chapter 3 promised you and left here. Move six is Chapter 1's pricing argument with a number behind it.
Only move five is new.
The check is not new either. Chapter 2 got there first. It argued that your buyer's vocabulary beats your own.
What is new is the price tag. Chapter 2 framed buyer vocabulary as something to add. This chapter prices what it costs when somebody removes it. On your invoice. In a rewrite you approved.
A sentence has two layers, and only one of them is safe to edit. One layer per gate, in the order Figure 5.0 puts them.
The claim layer is what your sentence asserts. How firmly, about whom, on what evidence, as of when.
Almost every controlled result in this chapter that came back positive lives there.
The vocabulary layer is the words you chose to say it in.
Almost every controlled result that came back negative lives there.
So edit the claim. Leave your words alone.
Which is the reverse of what almost every content service sells. So ask yours: which layer do you edit?
Rewriting vocabulary is the part that is easy to sell. It is also the part that shows up in a before-and-after screenshot. You have seen that deck.
- Name the subject inside the sentence. Not in the heading above it. Not in the paragraph before it. If the sentence opens with "It," "This" or "Our platform," the subject is sitting somewhere the reader may never receive.
- Delete the unbounded qualifier, or replace it with a number. "Usually," "typically," "most," "fast," "industry-leading." An unbounded qualifier is not a cautious claim. It is no claim, and a model that cannot state it cannot ground on it.
- Attach the condition to the number. A figure without its unit, its population and its date is a figure somebody else has to caveat for you. "6 working days" is weaker than "6 working days, median, for a 50-seat rollout."
- Date anything that can go stale, in the prose. Recency was one of only four factors clearing every model in the SIGIR study. Put it in the sentence, not only in the byline, because the byline may not travel with the sentence.
- Say who verified it, if anyone did. Evidence attached ran ahead of evidence absent in all six models, significantly in four. An audit, a test, a certification, a sample size. One clause is enough.
- Put a number on the price, or a range. "Price Not Mentioned" was another of those four gatekeepers, and it is the one most B2B pages fail. Chapter 1 already answered your legal team: bound the range, date the claim, attribute it to a named source.
Here is the same edit run on a whole sentence. And on the version an agency would ship instead.
The objection this chapter has to answer
"You told me not to upgrade my vocabulary, because sparse retrieval loses term overlap. Then you told me real engines run dense, where the cost is 0.95 rank positions. From a single-run study with no significance testing.
"Your headline finding is a rounding error on the retriever that actually serves my query."
That is the right objection. Sharper than the one about the vendors disagreeing.
Three answers. The direction is unanimous across three teams, and two of them never used BM25 at all. The 0.95 is an average, and the tactics being sold to you sit at the bad end of the spread. Not the middle.
And the size is not really the point. The point is that the trade is invisible. Whatever it costs, you are paying for the side of it that loses. No report you own separates the two.
Then the wider objection, because it is coming anyway. Two engines disagree. Neither ran an experiment. Four studies measure four things. The best of them prints numbers it tells you not to trust.
Why act on any of it? Table 5.3 shows you what the alternative is. Worse.
Take that objection at full strength too. This is the thinnest evidence base in the book, and it is not close. By some way.
Chapter 3 rests on vendor documentation you can verify with curl in an afternoon.
This chapter rests on four studies that disagree, and one style guide from an advertising blog.
So the method above is built to be cheap and to fail safe.
Naming your subject, bounding your numbers and dating your claims cost a writer twenty minutes. They improve the page for humans whatever the engines turn out to do. That is the hedge you want.
And the check costs nothing at all: it is an instruction to stop doing something. How often do you get one of those?
Nothing in Artifact 5.4 asks you to believe a single number in this chapter. That is deliberate. It is also how to treat any advice about a system nobody has documented.
What is left: six instructions
The engines have published almost nothing about writing. What little they have published contradicts itself. The studies measure prominence inside a context you already reached. All but one. That one measures whether you reach it. It says every rewrite costs you something.
Against all that, six things hold.
Name your subject. Bound your claims. Date them. Say who checked them. Put a number on the price. Do not upgrade your vocabulary.
Six instructions, and your writers can hold all six in their heads: subject, bound, date, verify, price, and leave the words alone.
That is the only reason any of them will survive a quarter.
The first five are unglamorous and free. The last one requires canceling something.
Which is why it is the one worth an argument in your next content review.
One loose end first, and Chapter 11 picks it up. Nothing in your reporting separates a rewrite that helped from one that quietly cost you.
That chapter exists for exactly this.
Now notice what the SIGIR study put at the top of its list. Above every wording factor it tested.
Topic mismatch. Not phrasing. Whether your page is about the thing at all.
Which is a question about what your page is about, and that is not a wording problem. It is a bigger one.
Chapter 6 is where it gets solved.
- Edit the claim, not the words. Then run the check: every content noun and verb you introduced, looked up in your query data.
- Date anything that can go stale, in the prose. Recency cleared all six models, and the byline may not travel with the sentence.
- Put a number on the price. A missing price cleared every model as a gatekeeper, and most B2B pages fail it.
- Ask any number for its outcome variable. Prominence inside a context you reached is not the same result as reaching it.
- Do not buy tone as a lever. Null in the paper that invented it, negative in two more, every positive after retrieval.
- Cancel the llms.txt ticket. Four studies, no effect. 97% of the files get zero requests, and Google has twice said no.
- Krishna Madhavan, Microsoft Advertising, "Optimizing Your Content for Inclusion in AI Search Answers," 8 October 2025. The four snippet-eligibility criteria, quoted in part. A Microsoft Advertising blog post, linked by the Bing Webmaster Tools blog as the deeper guidance, carrying no last-updated date.
- Bing Webmaster Guidelines, sections 10 and 15 to 18. Grounding and citation guidance, entity naming, single topic per URL, key information early. The section 10 instruction to use a "data-snippet" attribute has no corresponding documentation anywhere Microsoft publishes.
- Google Search Central, "Optimizing your website for generative AI features on Google Search," published 15 May 2026, updated 10 July 2026. Mythbusting entries on rewriting for AI, machine-readable files and structured data.
- Google Search Central, "Creating helpful, reliable, people-first content," updated 10 December 2025. The word-count question, verbatim.
- Perplexity, "Architecting and Evaluating an AI-First Search API," 25 September 2025, and "Search API: better extraction, dynamic benchmarks," 11 March 2026. Span labeling published, criteria never published.
- Anthropic, Claude Platform citations documentation. Sentence-level chunking, for developer-supplied documents only. OpenAI publishes nothing to website owners about text.
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, "GEO: Generative Engine Optimization," KDD 2024, arXiv:2311.09735. Position-adjusted word count, baseline 19.3, Quotation Addition 27.2, Easy-to-Understand 22.0, Authoritative 21.3. Sources fixed at the top five results, so retrieval is held constant. Authoritative rewriting scored 21.3 against the 19.3 baseline, and the paper's own gloss is that "we find no significant improvement, demonstrating that Generative Engines are already somewhat robust to such changes."
- Puerto, Gubri, Green, Oh and Yun, "C-SEO Bench: Does Conversational SEO Work?", NeurIPS 2025 Datasets and Benchmarks Track. Three of 54 cases significant. One of the two effective methods is a markdown summary rather than a rewrite, and the other is defined as all eight style transformations applied at once plus structural formatting.
- Kim, Jeong, Kim, Lee and Lee, "SAGEO Arena," KDD 2026, arXiv:2602.12187. Ten rewriting strategies. Body-text-only average rank change of 4.54 positions under BM25, 2.81 hybrid, 0.95 dense, from their Table 4. The AutoGEO-style strategy placed tenth of ten at 22.35 positions. No significance testing, no variance and no repeated runs are reported anywhere in the paper, and the pipeline is BM25 with a Qwen reranker rather than a commercial engine.
- Wu, Zhong, Kim and Xiong, "What Generative Search Engines Like and How to Optimize Web Content Cooperatively" (AutoGEO), ICLR 2026, arXiv:2510.11438. A 50.99% average gain over Fluency Optimization across three datasets, on the share-of-answer metric, over a fixed five-document candidate set. The version SAGEO Arena evaluated is a prompt reimplementation rather than the released system, which is why this chapter reports the comparison and not a verdict.
- Bagga, Farias, Korkotashvili, Peng and Wu, "E-GEO: A Testbed for Generative Engine Optimization in E-Commerce," arXiv:2511.20867, version 2, 14 July 2026. 13,747 queries. Technical and Unique in the negative bucket of fifteen hand-written heuristics, and Authoritative negative across all five re-rankers. Unique turns positive after the paper's own prompt optimization. Technical only reaches about zero, and stays negative on three of the five re-rankers, which is a result about prompts rather than about wording.
- Wan, Wallace and Klein, "What Evidence Do Language Models Find Convincing?", ACL 2024. ConflictingQA, 238 questions, 2,208 retrieved paragraphs, five models. Flesch-Kincaid and unique-token count show no correlation with judged convincingness. Reported as figures rather than coefficients, so quote the sentences and not a number.
- Van de Sande, Acar, van Woudenberg and Larson, Radboud University, "On Fact and Frequency," arXiv:2503.04271. Verified-false claims rewritten with uncertainty markers, meaning held fixed and every rewrite hand-checked, then fact-checked by three models. Hedging flipped 25% of classifications from false to not-false. Belief framing flipped 50 to 56%. A fact-checking task, not a citation task.
- AirOps, "Structuring content for LLMs." 12,000 URLs across 900 queries and 15 industries. 68.7% of ChatGPT-cited pages used sequential heading structure against 23.9% of Google's top results, which is a prevalence ratio between two corpora and not a measured citation lift.
- Microsoft, "Announcing Microsoft Web IQ," 2 June 2026. Passages and structured evidence objects, in a grounding infrastructure product rather than a consumer surface.
- Vishwakarma, Kumar and Jamidar, "What Gets Cited: Competitive GEO in AI Answer Engines," SIGIR 2026, arXiv:2605.25517. 252,000 trials, 18 factors, six models, two documents per trial. Outcome is which source receives the first citation. Bold denotes significance, ">10k" denotes quasi-separation rather than an effect size, and the asterisk and dagger denote degenerate Hessian and singular fit. No multiple-comparison correction is reported. The corpus is product review blogs, generated by another model.
- Evan Hall, Portent, "Featured snippet display lengths," 3 June 2021. 7,854 featured-snippet keywords, desktop only. It replicates an earlier Ghergich and SEMrush finding on word count, and reports the two to three sentence result separately. Between them they are the origin of both circulating rules.
- Jeremy Howard, Answer.AI, "The /llms.txt file," 3 September 2024. John Mueller on Reddit, 17 April 2025, and Gary Illyes at Search Central Deep Dive APAC, 23 July 2025. SE Ranking, roughly 300,000 domains, 7 November 2025. Server-log analysis across 137,000 domains, May 2026. Otterly.ai, 5 February 2026. Previsible, ten sites, 20 January 2026.
- SE Ranking, 129,000 domains and 216,524 pages across 20 niches, 24 November 2025. Question-style headings averaged 3.4 citations against 4.3 for straightforward headings. The same article reports a SHAP attribution on FAQ presence running the other way, and recommends question-based titles, reporting close to seven times the effect on smaller domains. Both figures are descriptive, and neither is an experiment.