Reddit's robots.txt has two directives. Both say no, to everyone. That is one of the two names Chapter 8 ended on.
Neither of these is a channel
Chapter 8 ended on two names. YouTube mentions topped the largest correlation table at 0.737. Reddit discussion was the strongest predictor on the one engine where anything predicted. Two names, both handed here.
It also said neither is a placement you buy, and handed the question here. So here is the answer, and it is not a posting strategy.
These two platforms are structurally different from each other, and both are different from the open web. One reaches engines through commercial contracts that can be canceled. The other lets anyone watch and almost nobody read.
Neither is a surface you optimize. Both behave like infrastructure somebody owns. So what do you do?
Which changes what you should do, and mostly it means doing less of what you were sold.
Reddit is not a channel. It is a contract.
Start with the file every crawler reads first.
# Welcome to Reddit's robots.txt
# Reddit believes in an open internet, but not the misuse of public content.
# See https://support.reddithelp.com/hc/en-us/articles/26410290525844-Public-Content-Policy Reddit's Public Content Policy for access and use restrictions to Reddit content.
# See https://www.reddit.com/r/reddit4researchers/ for details on how Reddit continues to support research and non-commercial use.
# policy: https://support.reddithelp.com/hc/en-us/articles/26410290525844-Public-Content-Policy
User-agent: *
Disallow: /
Eight lines, and five are comments. The two directives have no exceptions, no named agents, no allow list.
The comments point at two other doors: the Public Content Policy, and a subreddit for non-commercial research.
Both doors lead to one room. Have you read either?
The Public Content Policy says it in one line: you may use Reddit content for non-commercial uses, "but talk to us if you have commercial purposes in mind."
The unauthenticated JSON endpoint that used to serve any subreddit as structured data returns a 403 now. The Data API terms close the last door explicitly.
Their section 2.4 grants no rights to use content "for other purposes, such as for training a machine learning or AI model, without the express permission of rightsholders."
No ambiguity there.
Section 3.2 repeats it as a prohibition rather than an omission. So how does any Reddit content reach an engine?
By contract. A small number of them.
| Deal | Announced | What the announcement claims | Says "train"? |
|---|---|---|---|
| Google and Reddit | 22 February 2024 | Structured access to "display, train on, and otherwise use" Reddit content | Yes, explicitly |
| OpenAI and Reddit | 16 May 2024 | Data API access to "bring enhanced Reddit content to ChatGPT and new products" | No. Retrieval and display only |
Now price the whole arrangement, because it is smaller than the discourse suggests.
Reddit's Other revenue line is mostly data licensing. It ran at 43 million dollars in the second quarter of 2026. Against 762 million in advertising.
Roughly 5% of revenue: split across every licensee.
The largest of those contracts may be ending. The Google deal is reported at 60 million a year. In July 2026 Reddit was reported to be weighing whether to renew. Internally. Quietly.
The stock fell 9% that day. Which tells you how the market reads that channel.
So hold the whole picture. Reddit blocks every crawler. It forbids training through its API. It licenses to a handful of buyers, for about a twentieth of its revenue. That is the whole arrangement.
Is that a surface you optimize? No. It is a pipe two companies negotiate over. Either can hang up. We would not build on that.
Which is the first thing this chapter tells you not to build on.
Reddit is defending that pipe in court, which tells you how it sees the asset.
In October 2025 it sued Perplexity and three data-scraping companies. The complaint is worth reading for its method, not its law. The method is a trap.
Reddit says it planted a test post visible only to Google's crawler. A marked bill.
Their words: "Within hours, queries to Perplexity's answer engine produced the contents of that test post."
The claims are not copyright infringement. Three counts sit under the DMCA's anti-circumvention provisions, and three are state-law claims.
Then the ruling, on 31 July 2026, and it matters beyond this case.
The court let the core circumvention claims proceed. It held that a bot-detection system counts as "a technological measure that effectively controls access." Even where a human can still read the page.
Read what that establishes. Routing around a CAPTCHA to collect public content is now a claim that survives a motion to dismiss.
The trafficking count went. So did the two economic claims under state law, as preempted. The conspiracy count survived, and nothing is decided on the merits.
But if you buy data from anybody who scrapes, that ruling is your risk too. Do you know who your vendors buy from?
Reddit treats access to its content as a thing you buy or a thing you steal. There is no third category, and you are not in the first one.
A strange position for a platform built on public posting. We would plan around it.
One more thing about Reddit, and this one is useful. When it does get cited, it does not concentrate where you would guess.
A vendor study of 1,486 business software queries found 365 subreddits appearing across results. Only 96 reached an answer box. The top twenty absorbed half of all appearances.
Now the names. The most cited were narrow. A customer relationship management subreddit at 25 citations. A software-as-a-service one at 14. Self-hosting at 11.
Not the giant general subreddits: the small specific ones.
A second vendor found the same shape from the other end. The median cited comment: 38 upvotes. And 71% of cited threads sat in subreddits under half a million members. Small rooms. Cheap rooms.
Both numbers come from companies selling visibility software, so hold them loosely. But they agree with each other, and they point somewhere counterintuitive.
What gets surfaced is not the popular thread. It is the specific answer, in the narrow room, where somebody asked your exact question. One person, once.
A much smaller target than a Reddit strategy. And a cheaper one.
YouTube is not closed. It is asymmetric.
YouTube runs the opposite way, and the asymmetry is the finding.
Google's Gemini API documents it plainly. You pass a YouTube URL straight into a request. The model watches the video.
Their constraints: public videos only, ten per request, no length cap on the paid tier.
So Google's model can read any public video on the platform Google owns. So can you, through the same documented route, with an API key.
Which is where most write-ups of this stop, and it is the wrong place to stop. Two things complicate it.
First, that door is marked temporary. Google's own banner: the YouTube URL feature "is in preview and is available at no charge. Pricing and rate limits are likely to change."
Second, and this is the real asymmetry, watching is not the same as having the text.
The caption track is what every transcript-based tool needs. The Data API is one sentence long on who may download it.
The method "requires the user to have permission to edit the video."
Read what that closes. You can pull captions for one set of videos: your own.
Not your competitor's. Not the review ranking you fourth. Not the tutorial naming a rival in the title.
You can have a model watch all three. You cannot have the file.
Ask anyway: a 403. "The permissions associated with the request are not sufficient to download the caption track." Not yours to read.
So the practical consequence, stated plainly. Nobody outside Google can build a transcript corpus of videos they do not own.
Which is a narrower complaint than the field makes, and a more useful one. Any competitive YouTube number built on transcript text came from scraping, estimation or a vendor's private index.
Ask which one. Ask it now that a court has said routing around bot detection survives a motion to dismiss.
Worth knowing before you buy a dashboard built on one. Which one are you paying for?
What the rules actually say
Both platforms have brand participation rules. One set is enforceable law. Most GEO advice acts as though neither exists.
Start with the law, because the exposure is real and priced.
The Federal Trade Commission's rule on consumer reviews and testimonials took effect on 21 October 2024. It bans six things. Fake reviews. Buying reviews of any sentiment. Undisclosed insider reviews. Company-controlled review sites. Review suppression. Fake social-media indicators.
The civil penalty runs to one number: 53,088 dollars per violation.
In December 2025 the Commission sent warning letters to ten companies. Not determinations, but a signal about where attention is going.
Then the older instrument, and it is the one that names your exact situation.
"An online community has a section dedicated to discussions of robotic products. Community members ask and answer questions and otherwise exchange information and opinions about robotic products and developments. Unbeknownst to this community, an employee of a leading home robot manufacturer has been posting messages on the discussion board promoting the manufacturer's new product. Knowledge of this poster's employment likely would affect the weight or credibility of the endorsements. Therefore, the poster should clearly and conspicuously disclose their relationship to the manufacturer. To limit its own liability for such posts, the employer should engage in appropriate training of employees. To the extent that the employer has directed such endorsements or otherwise has reason to know about them, it should also be monitoring them and taking other steps to ensure compliance. (See § 255.1(d).) The disclosure requirements in this example would apply equally to employees posting their own reviews of the product on retail websites or review platforms."
Read the first two sentences again. An online community, a section for one product category, members asking and answering questions.
That is a subreddit, described by a regulator, before subreddits were the strategy.
The standard behind it is broader than most people think. Any connection that "might materially affect the weight or credibility of the endorsement," and that the audience would not expect, "must be disclosed clearly and conspicuously." Clearly. And conspicuously.
The duty does not stop there. Employers are expected to train staff and monitor compliance.
Now the useful half, because the rule is narrower than the panic around it.
| The action | Status | The condition that decides it |
|---|---|---|
| Employee posts about your product | Allowed | They disclose the relationship clearly and conspicuously |
| Incentivized customer reviews | Allowed | No requirement, express or implied, that the review be positive |
| Paid YouTube placement | Allowed | The creator ticks the paid promotion box, and you are jointly responsible for legal disclosure |
| Undisclosed employee posting | Banned | Named in the Endorsement Guides since before generative engines existed |
| Buying reviews of any sentiment | Banned | Rule effective 21 October 2024, up to 53,088 dollars per violation |
| Vote manipulation on Reddit | Banned | Reddit's guidance says asking people by message to vote "will result in a ban from the admins" |
Reddit's own rules are softer than the law. Worth reading for what they are.
The famous nine-to-one ratio, one self-promotional post in ten, lives in Reddiquette. That is informal guidance rather than the enforceable content policy.
So it is a norm the community polices, not a rule the company enforces. Which is harsher? The community, every time.
The company is now paying attention, for a new reason.
In July 2026 Reddit reported catching roughly 25,000 net new spammy posts and comments a day. It uses language models to find "coordinated patterns of fake behavior and artificial hype." Models catching models. Which side is your agency on?
Reddit's own post never says what the spammers are after. The trade coverage does.
EMARKETER, on 6 July 2026: as marketers optimize for AI answers instead of Google results, "a new spam economy is emerging."
So the motive is not Reddit traffic. It is getting cited.
So the tactic your agency may be pitching is already being run at scale. And caught at scale.
Though read Reddit's own enforcement numbers before you assume the risk is evenly spread.
In its most recent transparency report, covering late 2025, spam was 54% of administrator removals.
Go back to the January to June 2024 report, which broke the categories out further. Spam was 66.5% there. The category covering vote manipulation and artificial promotion: 1.8%.
Which is the honest picture: Bulk spam is the volume problem. Coordinated brand manipulation is a small labeled slice.
So the enforcement risk is real, and uneven. The legal risk in Table 9.3 does not care about distribution at all.
Four readers, four exposures
The legal exposure in Table 9.3 is not evenly distributed either. Where you sit decides which row you should worry about.
First, business software. Your category is where the concentration data came from. Those narrow subreddits are real rooms where buyers compare tools.
Your exposure is Example 8. A support engineer answering a thread is your best asset and your likeliest violation. The difference is one sentence.
Second, consumer and ecommerce. You carry the highest exposure in this chapter. Reviews are your channel. The review rule is the one with a number attached.
Audit how you solicit reviews before you audit anything else. Any incentive conditioned on sentiment, anywhere in your flow, is this week's job. Find it now.
Third, regulated categories. Health, finance, legal.
Your own regulator almost certainly has stricter endorsement rules than the Commission's. This chapter is the floor, not the ceiling. Ask your compliance team before step 5, not after.
Fourth, publishers and creators. You are on the other side of this. YouTube's declaration checkbox is your obligation, not the brand's.
The platform pushes legal disclosure to both of you. Which means a brand's failure to disclose becomes your problem too.
Traced to source
Fifth time. Chapters 5, 6, 7 and 8 each walked five claims back to origin. Same method here.
| The claim | Where it comes from | Does it hold? |
|---|---|---|
| "GPT was trained on Reddit" | GPT-2's WebText, built from 45 million links posted by Reddit users | Links Reddit users shared, and the model card says explicitly it is "not data taken directly from Reddit itself" |
| "Google reads YouTube transcripts for AI Overviews" | No origin document. Gemini can ingest a YouTube URL, which is a different product and a different claim | Neither doc uses the word transcript |
| "Reddit is 6.6% of Perplexity citations, so post there" | A vendor citation-share dataset, priced in Chapter 6 alongside Pew's independent panel | A share of citations is not a share you can enter |
| "The Google Reddit deal is worth 60 million a year" | A single Reuters report of February 2024, restated everywhere since | Neither company has published the figure |
| "Seed the discussion and the model follows" | No published before and after exists, on either platform | Untested. The field's survey of 45 studies reviewed no study of community seeding at all |
One query, walked
Take one query and follow it, because the abstractions above have a shape you can see. Somebody types what your buyers type: best tool for X, or alternatives to your competitor.
Ask what the engine has to work with. Four or five slots, not ten.
What fills them? On the independent numbers: mostly official sites, news and vertical publications, at roughly four fifths of citations.
Then the rest. One slot might be a video. One might be a thread in a subreddit with 40,000 members. Somebody asked your exact question there two years ago and got eleven replies.
Now ask what you could have done about each.
The official and vertical sources: that is Part II and ordinary coverage. The video: yours is measurable, everyone else's is not.
The thread: you could have answered it, with disclosure, when it was posted. You cannot answer it now in a way that reads as anything but late.
Which is the honest shape of this work. It is not a campaign. It is a presence you either had or did not have when somebody asked.
The rest of the corpus
Two platforms is not the whole off-site world, and Chapter 3 named the others. G2. Comparison pages. Your competitors' alternatives-to posts.
So take the verdict on those, briefly, because it differs from everything above.
They are open. No robots.txt wall, no licensing contract, no edit-permission gate.
Which means an engine can reach them the ordinary way, and so can you. You can read your category's comparison pages today. Free, in a browser.
Review sites sit between. Most run their own terms. Most have a front door. Most are explicit that you may solicit reviews.
Which puts them under Table 9.3 rather than under this chapter's platform mechanics. Solicit freely. Never condition on sentiment.
Then the alternatives-to page written by a competitor. That one is not a corpus problem at all.
It is a page describing you, that you did not write, ranking for your name. The response is a better page of your own. Which is Part II.
So the whole off-site corpus is three tiers: Open pages you can read and answer. Review platforms with rules and a front door. Then two large platforms that are closed or asymmetric.
Most budgets are aimed at the third tier. Most return sits in tier one. Where is yours?
What these platforms are actually good for
This chapter has been negative for six sections. Here is the honest case on the other side.
Three things are genuinely worth having, and none of them is a visibility tactic.
The first is the only unfiltered read you will get on your product. Step 2 of the artifact is not a marketing exercise.
Somebody described your onboarding to a stranger, unprompted, in their own words. No research budget buys that sentence.
The second is measurement. YouTube is the one platform where your own footprint is fully readable by you. You can pull captions for your own videos.
Which makes your own channel the cheapest place to test one thing. Is the language you use the language buyers repeat? Chapter 5 gave you the method.
The third is durability. Answering a real question in public outlasts anything else in this book. It sits there, dated, attached to your name, useful to the next person with the same problem. For years.
If a model reads it, good. If not, a customer did.
That is the whole case, and notice it does not depend on anything an engine does.
The direction of travel
Everything above describes the platforms as they are. Now the part that should change your allocation. The trend runs against Chapter 8.
Google said so itself, on the record, in May 2024.
The sentence came after the week AI Overviews suggested glue on pizza. "We updated our systems to limit the use of user-generated content in responses that could offer misleading advice." Limit. Their word, not ours.
The same post calls forums "a great source of authentic, first-hand information" in the sentence before. Both halves are true, and only one of them got engineered.
Then the measurement, and here you have to be careful.
Conductor tracked Reddit's citation share falling through late 2025 and into 2026. It puts the share at 2.02% in October and 1.01% by January, measured across engines rather than on any single one.
Take one thing from that: the direction, not the decimals.
Because the same datasets cannot agree what the level is. Chapter 6 carried 1.8% for ChatGPT across all queries. Promptwatch had ChatGPT Search steady at 3.83% in midsummer. MaxAEO reports 8.1% on ChatGPT, on recommendation-intent prompts only.
Different denominators, not a trend. Anyone drawing a line through them is drawing it through three different questions. Not one trend.
Which is the whole problem with this evidence base, stated once. Almost every number you will be quoted here comes from a company selling visibility software. Including ours. Weigh us accordingly.
Then, six days before this chapter was finished, the argument made itself.
That Promptwatch 3.83% held from 18 July through 7 August. On 14 August it broke. The average from 14 to 17 August was 0.52%, an 86% fall in four days.
Google's surfaces did not do that. AI Overviews drifted down 11% across the same weeks, AI Mode about 30%. Gradual, both.
Nobody has said why, and the vendor calls its own finding provisional. It cannot rule out a fault in its own collection.
So do not build on it. Build on the shape of it.
A surface reached by contract can change overnight, with no announcement, and you find out from a dashboard. That is what not a channel means in practice. Would you have noticed?
So weigh the independent measurement separately, and it is more sober.
An academic study of 602 controlled prompts and 21,143 citations found YouTube the most cited single domain. Reddit came third. Both real, and both small: YouTube around 2.6% of citations, Reddit around 1.5%.
A second, of 366,087 citations, found social media platforms at 10% of all citations put together.
Which is a real channel. It is not the channel a 0.737 correlation makes you picture.
One more finding, and it explains why every share in this chapter looks small. A study of 55,936 queries found generative engines return 4.3 URLs per response, against 10 for traditional search.
Fewer than half the slots. So every domain competes in a smaller room. Single digits mean more here.
The same study found something else. Roughly 37% of the domains generative engines cite never appear in traditional search at all. Which cuts both ways for you.
Fewer seats. Different guest list.
- Count what already exists, without touching anything. Search Reddit for your brand name and your category, and count threads you did not start in the last 12 months. Then YouTube, for videos naming you that you did not pay for. Two numbers. Most companies have never written them down.
- Read the ten most recent, all the way through. Not the sentiment score a tool gives you. The actual complaint, in the actual words, because this is the text a model may be summarizing when somebody asks about you.
- Separate what is wrong from what is unflattering. Factually wrong is a correction you can request through normal channels, disclosed. Unflattering and true is product feedback, and no posting strategy fixes it.
- Write the disclosure line before anybody posts anything. One sentence naming who you work for, agreed with whoever owns legal risk. The Endorsement Guides require training and monitoring, not just the disclosure itself, and having the sentence pre-approved is how a rule becomes a habit.
- Answer only where you are already named. Threads that mention you, questions about your category that you can answer without a link. Never a thread you seeded, never an account that hides the relationship, and never a review you incentivized on condition of sentiment.
- Re-run the two counts in ninety days, and do nothing else. If they moved, note by how much and whether you can attribute it. If a vendor tells you those counts drove an AI visibility change, ask which study established that link, and read Table 9.4 row five while they look.
What this costs
Price it before you fund it, the same way Chapters 7 and 8 did. Community work has three costs, and only one is obvious.
The visible one is time, and it is expensive time. Answering well in a technical subreddit takes somebody who knows the product. Really knows it. Who is that, here?
That is a support engineer or a founder. Not a coordinator, and not an agency writer with a brief.
The second is compliance overhead. The Endorsement Guides expect training and monitoring, not just a disclosure line. So any program with more than two participants needs a record of who was told what. Write it down.
Small, but real. It appears in nobody's proposal.
The third is the one that ends most programs. Public participation means public criticism. In a thread you cannot edit, attached to your name, permanently.
Somebody has to be willing to answer the hard reply. If nobody senior is, the program stops after the first bad thread. You have paid for the setup.
Against that, the honest return. A count you can measure. A corpus you can read. And no published evidence linking either to an AI answer.
Which is a smaller promise than the category makes. Fund it accordingly, or do not fund it.
If you are already doing it
Some readers will get here having already run undisclosed accounts. So take the remediation path rather than the lecture.
Stop first, and stop today. Every further post compounds an exposure that is priced per violation.
Then find out what exists. Four questions: which accounts, which threads, whose idea, and was anyone paid?
Take that to whoever owns legal risk before you take it anywhere else. This chapter is not legal advice and the next step genuinely is.
Do not mass delete. Deleting is what investigators look for. Reddit's license terms already bar partners from using deleted content. So removal does not clean the corpus you were worried about.
Then rebuild it, disclosed. In most cases the posts were useful and only the attribution was missing. A cheaper repair than people fear.
The objections this chapter has to answer
"Reddit threads outrank my site for my own category terms. I can see it. You are telling me not to compete there?"
No. We are telling you what competing there consists of: answering, disclosed.
Answer where you are named, with disclosure, and that is competing. What this chapter rules out is the version with no disclosure on it.
A second, and it is the strongest commercial objection.
"Everyone in my category is doing it. If undisclosed seeding works and nobody enforces the rule, I am choosing to lose."
Two answers, and neither is moral:
The first is that the enforcement picture changed. Reddit is running language models against exactly this pattern, at roughly 25,000 items a day. And the Commission sent its first warning letters in December 2025.
The second is that nobody has shown it works. Not one published study, on either platform, has added activity and measured an AI visibility change. Not one.
So the trade is a rule with a per-violation penalty, against a benefit nobody has demonstrated.
A third objection, from a reader who took Chapter 8 seriously.
"You just spent a chapter telling me off-site mentions are what accumulate. Now you are telling me the two biggest sources are closed. What is left?"
Fair, and the answer is in Chapter 8's own table. Branded web mentions sat at 0.664, just under YouTube. Those are ordinary earned coverage.
The open web is not closed. It is only slower, and it was always the larger share.
Then a fourth, which cuts at our sourcing. "Half your citation numbers come from vendors selling visibility tools, including the ones you quote approvingly."
Correct, and we flagged it in the section itself. Hold the independent measurements instead. 602 prompts and 21,143 citations. Then 366,087 citations across twelve models.
Both are low single digits. Ten percent for social overall. Where vendor data and academic data disagree, we printed both. And told you which is which.
Then the one we raise ourselves.
One caveat on method, and it is a strange one. Reddit's block on that file keys to the agent string, not to whether you are a crawler.
Testing on 20 August 2026, curl got the file. So did an invented bot name we made up. GPTBot, Googlebot and bingbot each got a 403. Why those three?
We do not know. What you see depends on how you ask, so we fetched it with a browser agent and checked the Data API terms separately. They say the same thing.
The disclosure this chapter owes you. We sell GEO services. Community and coverage work is the largest line we could bill you for.
This chapter narrows that line to disclosed participation and a counting exercise. Notice what it does not do. It does not sell you a subreddit strategy.
What would change our mind
This chapter says do less. So it should say what evidence would make us say otherwise.
Three things, and any one would move us:
The first is a real before and after. Matched brands. One participating with disclosure, one not. Appearance rate on Chapter 8's schedule.
Nobody has run it. It is not expensive, and the field that sells this has had three years.
The second: an engine saying anything. Any documentation describing how community content is weighted would replace a chapter of inference with a paragraph of fact.
Google's only on-record statement points the other way, at limiting it.
The third is Reddit publishing what its licensees do. Volumes, a partner list, a split between training and retrieval.
Right now the company at the center of this names categories of buyer and nothing else. Which is its right, and it leaves everyone downstream guessing.
Until one of those three arrives, the honest position is the one in Artifact 9.5. Count, participate where you are named, disclose, and do not pay for a causal story nobody has told.
What is left
So the ledger for these two platforms, in three lines:
Reddit reaches engines through revocable contracts worth about a twentieth of its revenue. It blocks everything else.
YouTube can be read completely by the company that owns it and only partially by you.
Then the independent measurement, which puts both in single digits of citations. In a year when the largest engine said, on the record, that it limited exactly this content.
None of which makes them worthless. All of which makes them one thing: a smaller line item than the pitch.
Which is the sentence to take into the budget meeting. Smaller, not zero. There is a difference, and vendors on both sides of this argument keep collapsing it.
What survives is unglamorous and it is the same thing Chapter 8 landed on. Be genuinely discussed, in places you do not control, by people who are not you.
There is no faster version. Every faster version in this chapter is either blocked, undocumented or against a rule with a number attached.
Which answers what Part III set out to ask. You now know what the web around you is made of. And which parts you can reach.
One question is left in this part, and it is narrower and more fixable. Before any of this matters, a machine has to be able to read your page at all.
Chapter 10 is about what a crawler can actually execute. Three earlier chapters have been sending debts to it.
- Stop treating Reddit as a channel. It blocks every crawler and forbids training through its API. Engines reach it by contract, and one is reportedly up for renewal.
- Accept that YouTube is unauditable. Any developer can have a model watch a public video. Caption download needs edit permission, so competitive numbers come from somewhere else.
- Write the disclosure sentence first. An employee posting without naming their employer is Example 8 in the Endorsement Guides, and it describes a discussion board.
- Incentivize reviews, but never sentiment. A discount for reviews is legal. A discount for a good review sits inside a rule at 53,088 dollars per violation.
- Answer where you are named. Do not seed. Reddit catches roughly 25,000 spam items a day, and trade coverage puts the motive at AI citation rather than Reddit traffic.
- Count threads and videos, not visibility. Two numbers, re-run in ninety days. Nobody has published a link between those counts and an AI answer.
- Reddit, robots.txt, fetched 20 August 2026. The complete file is eight lines and five of them are comments, pointing at the Public Content Policy and at r/reddit4researchers. The two directives are "User-agent: *" and "Disallow: /", with no allow list and no named agents. Reddit's response to a request for this file keys to the user agent string rather than to whether the requester is a crawler. Tested on 20 August 2026 it returned 200 to curl, python-requests, ClaudeBot, PerplexityBot, CCBot, AhrefsBot, Bytespider and an invented agent name, and 403 to GPTBot, Googlebot, bingbot, an empty agent string and a bare "Mozilla/5.0".
- Reddit, Data API Terms, retrieved 20 August 2026. Section 2.4 states that no rights are granted "including any right to use User Content for other purposes, such as for training a machine learning or AI model, without the express permission of rightsholders in the applicable User Content." Section 3.2 lists the same as a restriction. Section 3.1 requires a separate agreement for commercial use.
- Reddit, Public Content Policy, updated 29 May 2025. "You can use Reddit content for non-commercial uses, such as learning and community, but talk to us if you have commercial purposes in mind." Licensee categories named are brand monitoring companies, language model developers and academic researchers. License terms prohibit use of deleted content, sensitive demographic profiling, government surveillance and deceptive applications.
- Google, "An expanded partnership with Reddit," 22 February 2024. The Data API gives Google "efficient and structured access to fresher information" to "better understand Reddit content and display, train on, and otherwise use it."
- OpenAI, "OpenAI and Reddit Partnership," 16 May 2024. Data API access to "bring enhanced Reddit content to ChatGPT and new products," described as "real-time, structured, and unique content." The announcement does not use the word train. OpenAI also became a Reddit advertising partner.
- Reddit, Inc., "Reddit Reports Second Quarter 2026 Results," 30 July 2026. Total revenue 805 million dollars, up 61% year over year. Advertising revenue 762 million, up 64%. Other revenue, which is primarily data licensing, 43 million, up 24%. The 60 million dollar annual figure for the Google agreement originates in a Reuters report of February 2024 and has not been confirmed by either company.
- Wall Street Journal reporting as summarized 22 July 2026. Reddit had internally discussed ending Google's access for AI training as the agreement neared expiry, and the stock fell 9% on the report. As of 20 August 2026 no renewal or termination has been publicly confirmed.
- Reddit, Inc. v. SerpApi LLC, Perplexity AI, Inc., Oxylabs UAB and AWMProxy, No. 1:25-cv-08736-PAE, Southern District of New York, complaint filed 22 October 2025. Six counts: three under the DMCA anti-circumvention provisions at section 1201, plus New York unfair competition, unjust enrichment and civil conspiracy. There is no copyright infringement count. The complaint describes a test post "the equivalent of a digital marked bill" crawlable only by Google, and alleges that "Within hours, queries to Perplexity's answer engine produced the contents of that test post." On 31 July 2026 Judge Engelmayer "predominantly denies the motions to dismiss," sustaining the section 1201(a)(1)(A) claims against both defendants, the section 1201(a)(2) claim against SerpApi, and the New York civil conspiracy claim against both, while dismissing the section 1201(b) trafficking claim against SerpApi and the unjust enrichment and unfair competition claims as preempted by the Copyright Act. The holding that matters beyond the case is that a bot-detection system qualifies as a technological measure that effectively controls access even where the same content remains readable by a human. The conspiracy claim survives only as predicated on the DMCA violation, the court having pruned its two preempted objects. Opinion and Order, docket 104, 25 Civ. 8736 (PAE).
- Zhang, Ye, Peng, Garimella and Tyson, "Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines," arXiv:2512.09483, 10 December 2025. 55,936 queries across six generative engines and two traditional engines. Generative engines "return far fewer URLs per response (4.3 vs. 10 on average)," and 37% of the domains they cite are absent from traditional engine results.
- EMGI Group, Reddit citation analysis for business software queries, published May 2026. 1,486 buying queries across 18 categories. 365 distinct subreddits appeared across results and 96 were cited inside AI Overview answer boxes, with r/crm at 25 direct citations, r/saas at 14 and r/selfhosted at 11. The top twenty subreddits absorb half of all Reddit appearances. EMGI sells link building and AI visibility services, and has a direct commercial interest in the conclusion that community mentions drive AI visibility.
- MaxAEO, Reddit and ChatGPT citation study, first quarter 2026, drawn from roughly 25,000 cited Reddit URLs. The median cited comment had 38 upvotes, and 71% of cited threads sat in subreddits under 500,000 members. A separate MaxAEO analysis of 1.21 million citations behind recommendation-intent prompts, from 3,200 prompts weekly across twelve weeks in the first quarter of 2026, put Reddit at "8.1% on ChatGPT" and 17.9% on Perplexity. MaxAEO sells AI visibility tracking and optimization.
- Google, Gemini API video understanding documentation, retrieved 20 August 2026. "You can pass YouTube URLs directly to Gemini API as part of your request." Public videos only. Up to ten videos per request on current models, with no length limit on the paid tier and eight hours a day on the free one. The page carries a banner stating that the YouTube URL feature "is in preview and is available at no charge" and that "Pricing and rate limits are likely to change." The feature is available to any developer with an API key, not only to Google. The Vertex AI equivalent permits one YouTube URL per request. Neither document uses the word transcript, and neither states that YouTube feeds AI Overviews or Search grounding.
- Google, YouTube Data API, captions.download reference, retrieved 20 August 2026. The method "requires the user to have permission to edit the video," and an unauthorized request returns a 403 with the message that "The permissions associated with the request are not sufficient to download the caption track." The API revision history records no change to caption access in 2025 or 2026.
- Federal Trade Commission, Rule on the Use of Consumer Reviews and Testimonials, 16 CFR Part 465. Announced 14 August 2024 on a 5 to 0 vote, effective 21 October 2024. Prohibits fake or false reviews, buying positive or negative reviews, undisclosed insider reviews, company-controlled review sites, review suppression, and misuse of fake indicators of social media influence.
- Federal Trade Commission, 22 December 2025. Warning letters to ten companies regarding possible violations, including conditioning compensation on a particular sentiment and failing to disclose insider reviews. Civil penalties of up to 53,088 dollars per violation, a figure adjusted annually for inflation. The Commission stated the letters are "not formal determinations that the recipients have violated" the rule.
- Federal Trade Commission, Consumer Reviews and Testimonials Rule questions and answers. "The rule does not prohibit giving incentives for reviews, as long as there isn't an express or implied requirement that the reviews have to express a particular sentiment." Employees may review a product if they "clearly and conspicuously disclose their relationship to your business."
- 16 CFR 255.5, Endorsement Guides, disclosure of material connections, current as of 20 August 2026. The standard: where a connection "might materially affect the weight or credibility of the endorsement, and that connection is not reasonably expected by the audience, such connection must be disclosed clearly and conspicuously." Example 8 describes an online community with a section for discussing robotic products and an employee of a home robot manufacturer posting promotional messages there. Employers are expected to train employees and monitor compliance.
- YouTube Help, paid product placements, sponsorships and endorsements, retrieved 20 August 2026. Definitions of paid product placement, endorsement and sponsorship. Creators must declare paid promotion in Studio settings, and "You and the brands you work with are responsible for understanding and complying with local and legal obligations to disclose Paid Promotion in their content."
- Reddit, Reddiquette, retrieved 20 August 2026. Self promotion guidance that only one in ten submissions should be your own content. On vote manipulation, sending messages asking people to vote for your submission "will result in a ban from the admins." Reddiquette is informal guidance rather than the enforceable Content Policy.
- Reddit, Inc., "How We're Keeping Reddit Real and Safe in the AI Era," 6 July 2026. Reddit reports "Catching ~25K net new spammy posts and comments a day" and reducing spam exposure by roughly 20% from January to March 2026 against the prior three months, using language models to "catch the highly subtle, coordinated patterns of fake behavior and artificial hype that older systems once missed." Reddit's own post does not state a motive for the wave.
- EMARKETER, "Reddit's GEO crackdown could raise the stakes for AI visibility strategies," 6 July 2026. "As marketers start to optimize for genAI answers rather than Google Search results, a new spam economy is emerging." This is the attribution of motive that Reddit's own announcement does not make.
- OpenAI, GPT-2 model card, 2019. WebText is "the text contents of 45 million links posted by users of the Reddit social network," and the card states that WebText "does not consist of data taken directly from Reddit itself." Separately, the Dolma corpus cited in Chapter 8 contains 89 billion tokens of Reddit content directly, which is a different dataset.
- Google, "AI Overviews: About last week," 30 May 2024. "Forums are often a great source of authentic, first-hand information, but in some cases can lead to less-than-helpful advice, like using glue to get cheese to stick to pizza." Among the changes made: "We updated our systems to limit the use of user-generated content in responses that could offer misleading advice."
- Zhang, He and Yao, "From Citation Selection to Citation Absorption," arXiv:2604.25707, 29 April 2026. 602 controlled prompts and 21,143 valid citations across ChatGPT, Google and Perplexity. YouTube is the most cited single domain at 560 citations, Wikipedia second at 352, Reddit third at 315, which is approximately 2.6% and 1.5% of citations respectively. Official, news and vertical sources account for between 79.12% and 87.52% of citations.
- Yang, "News Source Citing Patterns in AI Search Systems," arXiv:2507.05301, 7 July 2025. 24,069 conversations and 366,087 citations across twelve models. Social media platforms account for 10% of citations. Low credibility sources are rarely cited. The author notes that news is only 9% of citations and that the remaining 91% deserve equal scrutiny.
- Conductor, Reddit citation analysis, covering October 2025 to January 2026. Reddit's citation share reported falling from 2.02% in October 2025 to 1.01% by January 2026, across 238,212 prompts where Reddit appeared. Conductor describes the dataset as "AI citation data across LLMs where Reddit was included as a cited source" and does not name the engines, so this is not a ChatGPT figure. Conductor sells enterprise search and answer engine optimization software. Other vendors report much higher baselines for adjacent metrics, which is why this chapter takes the direction and not the decimals.
- Promptwatch, Reddit citation tracking for ChatGPT Search, Google AI Overviews and Google AI Mode, 7 July to 17 August 2026, posted 18 August 2026 and reported by Search Engine Land on 19 August 2026. "From July 18 through August 7 it held a steady 3.83% average share of ChatGPT citations," and "the August 14-17 average of 0.52% is a 86.4% relative drop." Promptwatch also records a smaller step down on 8 August, from the high 3% range into the mid 2% range, so the fall is best described as four days rather than one. Over the same period AI Overviews moved from 2.37% to 2.10%, an 11.3% relative decline, and AI Mode from 2.22% to 1.54%, a 30.5% decline. Promptwatch writes that "The chart shows when each change happened, not why," that "a data-collection issue cannot be ruled out, so treat the size of the drop as provisional." No cause has been confirmed by OpenAI or Reddit. Promptwatch sells AI visibility monitoring.
- Martinez, "Optimizing Visibility in Generative Engines: A Critical Survey," arXiv:2607.14035, 15 July 2026. Reviews 45 studies. No reviewed technique "shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior." The survey also states that it does not assign non-peer-reviewed measurement studies the same evidentiary weight as published work, and treats third party mentions as an observational hypothesis rather than an established effect.
- Reddit Transparency Reports. The report covering July to December 2025 states that the majority of administrator removals of posts and comments were for spam, at 54%. The January to June 2024 report, which is the most recent to publish the finer category breakout, records 162,135,309 pieces of content removed, with spam at 66.5% of administrator removals and content manipulation, covering vote manipulation and artificial promotion, at 1.8%.