Two AI companies now pay Reddit directly for its content. But the licensing deals only explain part of why engines keep pulling from Reddit threads - the bigger reason is what those threads contain that brand websites structurally can't.
In February 2024, on the same day it filed for its IPO, Reddit signed a data-licensing deal with Google reported at roughly $60 million a year. It gives Google real-time access to Reddit's user-generated content for AI training and retrieval, and gave Reddit access to Google's Vertex AI platform in return. That deal is now in renewal talks - Reddit stock dropped on a July 2026 report that the company might tighten or withhold Google's access rather than auto-renew on the same terms, a sign Reddit thinks its data is worth more now than it was in 2024.
In May 2024, OpenAI announced its own Reddit partnership: API access to Reddit's real-time content, plus AI-powered features built into Reddit itself and an advertising relationship. Financial terms were never disclosed. Both deals mean Reddit content sits inside the retrieval pipelines of two of the largest AI answer providers by design, not by accident.
The licensing explains access. It doesn't fully explain why engines keep choosing to surface Reddit for a specific class of question. For informational or navigational queries, a brand's own product page is a fine source. For comparative and experience questions - "is this worth it," "what do people actually think," "best X for Y" - an engine retrieving from your marketing copy is retrieving from an interested party. It reads as promotional, not evidentiary.
Reddit threads are structurally the opposite: unaffiliated authors, visible disagreement, dated posts, replies that correct or contradict the original claim. That's what "first-hand experience" looks like to a system trying to answer honestly. No brand site can replicate that about itself, because the moment it's your own site making the comparison, it stops being independent evidence.
This is the part worth being honest about the limits of: precise citation-share numbers for Reddit vary by tracker, aren't independently audited, and shift as engines update retrieval. What's observable and consistent across public reporting:
The pattern that holds across all four: comparative and "is it legit" questions pull Reddit more than factual or navigational ones, regardless of engine.
What gets sold as a Reddit strategy is often a posting calendar: seed threads, plant comments, manufacture the appearance of consensus. What actually works is closer to what you'd do with any word-of-mouth channel you don't control - show up honestly where you have standing, be worth recommending, and measure what's already being said instead of trying to author it. The same discipline applies to other channels engines pull from that brands can't write for themselves; we cover the visual-platform version of this in does AI read Instagram, and the competitive angle in why ChatGPT recommends your competitor.
We ask AI engines the real comparative and "is it legit" questions buyers actually ask about a category, across ChatGPT, Claude, Gemini, and Perplexity, and read exactly what comes back - including when the answer is built from a Reddit thread rather than a brand's own site. That tells you whether the sentiment already circulating about you is helping or hurting before you spend a dollar trying to change it, and every count re-derives to a specific question and answer.