Platform Watch
  • GOOGL354.97+1.6%
  • MSFT501.61+1.6%
  • META741.25+11.4%
  • RDDT158.73+5.2%
  • HUBS213.86-1.0%
  • ADBE249.52+0.2%
  • CRM236.42-0.6%
AI Discovery

Reddit's AI Search Favors Formal, High-Vote Comments Over Personal Experience, Audit Finds

A University of Illinois Urbana-Champaign preprint studying 30,000 Reddit AI search answers found that already-upvoted, formally written comments were far more likely to be selected than comments rich in personal experience—and that most first-person language was stripped in the final answers.

Reddit's AI search feature surfaces formally written, already-popular comments far more often than comments that reflect personal experience, according to a preprint from researchers at the University of Illinois Urbana-Champaign. The study has not yet undergone peer review.

The research team ran 10,000 questions through the feature three times, generating 30,000 answers, which were then traced back through 14.68 million comments. The questions came from 20 advice and support subreddits, meaning the findings apply specifically to that segment of the platform.

Formality and Vote Score Were the Strongest Predictors

Among language signals, formality was the strongest predictor of selection. A comment that was one standard deviation more formal had 49% higher odds of being chosen, with an odds ratio of 1.488. Comments using directive language such as "should" and "must" also showed elevated odds (1.070). By contrast, comments with markers of personal experience had lower odds of selection (0.789), as did supportive language (0.924).

The researchers noted that formality correlated with higher community voting, while experiential voice correlated with lower voting. When controlling for a comment's score, position, and age, the associations narrowed—to 1.213 for formality and 0.860 for experiential voice.

Vote ranking within a thread showed the largest single association in the selection model. A one-standard-deviation increase in vote rank multiplied the odds of selection by 2.88. The median selected comment ranked at the 91st percentile for score in its thread, versus the 45th percentile for comments that were not chosen. Comments with a score of zero or below made up just 0.53% of selected comments, compared with 5.1% of all comments collected.

Additional structural factors also mattered: a one-standard-deviation increase in comment length produced an odds ratio of 1.79, and comments containing an external link had an odds ratio of 2.25. Among selected comments, 92% were direct replies to the original post. Selected comments appeared a median of 1.2 hours after the original post, while non-selected comments appeared after a median of 5.9 hours.

First-Person Language Nearly Vanished in Final Answers

In an analysis of 1,000 queries, researchers compared the quoted source comments with the AI-generated answers built around them. First-person terms such as "I" and "my" fell from 3.3% of words in the quoted comments to just 0.06% in the answers—a transformation the authors say converts personal testimony into generalized advice.

The same queries were also run through GPT-4o-mini and GPT-5 via OpenAI's API with web search enabled. Both models used less first-person language than the original Reddit comments, though Reddit's own responses showed the largest reduction. Out of 4,932 citations across both models and configurations, only 18 referenced Reddit.

How the Study Was Conducted

Researchers used a large language model to rewrite real posts from the 20 subreddits into short search queries. Ten subreddits were large communities, including r/personalfinance and r/AskDocs; ten were smaller, such as r/UKJobs and r/AusLegal. The three collection runs were spaced roughly five hours apart to minimize the effect of content changes on Reddit itself. A typical answer drew from about seven different subreddits.

The feature studied was originally launched as "Reddit Answers" in December 2024. As of a May 26, 2026 update, Reddit merged it into its main search experience. The questions used in the study were drawn from posts dated through July 2026, meaning the runs occurred after that integration.

Implications for Audience Research

The findings carry direct relevance for marketers and researchers who rely on Reddit for consumer insights. Reddit CEO Steve Huffman had argued to investors in February that Reddit excels at questions where "the answer actually is multiple perspectives from lots of people." This audit measured which of those perspectives actually reach the AI-generated answer—and found that the most experiential voices are disproportionately filtered out.

The authors caution that the study is observational. As they write, results "should not be interpreted causally." They also note that the selection patterns observed may not extend to other types of communities or other AI search systems.

Prepared with AI assistance and reviewed by the editorial team.

Sources

Continue reading

More in AI Discovery →
AI Discovery

NYT Filings Reveal OpenAI Knew of Paywall Workaround; Microsoft Executive Called AI Training 'Astonishing Theft'

Newly unsealed court filings in The New York Times' copyright lawsuit against OpenAI and Microsoft allege that senior figures at both companies were aware of efforts to access paywalled content and use copyrighted material to train AI models. Internal documents cite a Microsoft executive describing the practice as an "astonishing theft of unprecedented proportions" and internal data suggesting AI-powered search could reduce publisher traffic by up to 93%.

3 min read
AI Discovery

AI Search Is Reshaping Retail: What Brands Must Do Now

Conversational AI tools are fragmenting the traditional retail search journey, forcing brands to rethink how they structure product data, create content, and measure visibility. Experts say success now hinges on verified product knowledge graphs, schema markup, and cross-functional teams — not keyword density and backlinks.

7 min read
AI Discovery

Google DeepMind's Dream-RSI Lets AI Agents 'Dream' Their Way to Cheaper, Faster Search

Google and DeepMind researchers have introduced Dream-RSI, a technique that allows AI agents to test alternative search strategies by replaying recorded past attempts rather than running costly new computations. Benchmarked on Gemini 3.1 Pro and Gemini 3.7 Flash across eight tasks, the method matched or outperformed baselines while sharply reducing the number of live attempts required.

3 min read