Platform Watch
  • GOOGL354.97+1.6%
  • MSFT501.61+1.6%
  • META741.25+11.4%
  • RDDT158.73+5.2%
  • HUBS213.86-1.0%
  • ADBE249.52+0.2%
  • CRM236.42-0.6%
AI Discovery

AI Citation Study: Source Order Matters Less Than Raw Data Suggests

A new arXiv preprint testing GPT-5.4's citation behavior found that while top-ranked sources are cited far more often in raw data, deliberately swapping source order produced a much smaller and statistically insignificant effect — suggesting observed position gaps largely reflect content quality rather than rank alone. Structured formatting showed a modest but inconclusive citation lift, and model randomness caused citation decisions to flip in roughly one in seven reruns.

A new preprint posted to arXiv on September 14, 2026, challenges a common assumption in AI search optimization: that source ranking is a primary driver of AI citation. The paper, authored by researchers Sriram Selvam and Anneswa Ghosh, found that when matched sources were deliberately reordered, the citation impact was far smaller than raw position data would imply — and measured zero in a follow-up test.

How the Experiment Was Designed

The researchers prompted a GPT-5.4 search agent — using Exa as its search provider — to answer 130 common questions via independent web searches, recording all messages and search results from the 129 questions it addressed. They then identified pairs of pages that appeared in the same search results and were independently screened as supporting the same fact, leaving 113 valid pairs. A blinded human review later confirmed 103 of those pairs as genuine matches.

Each saved conversation was replayed four ways, varying whether a given page appeared above or below its counterpart and whether its text was formatted as plain paragraphs or rewritten with headings, lists, or a table. Crucially, the study used AI-generated rewrites — nearly all produced by Grok 4.3, with GPT-5.4 as a fallback for one pair — rather than live webpage edits. The paper notes this means the experiment compares two rewrites rather than isolating formatting differences in isolation.

The Position Gap: Bigger in Raw Numbers, Smaller Under Controlled Conditions

In raw terms, the positional gap looked substantial. Pages in the first position of a search call were cited 85.1% of the time, compared with 42.8% for pages in the fifth position — a difference of 42.3 percentage points. However, the study emphasizes that search providers typically surface more relevant pages at the top, meaning the raw gap reflects both position and underlying content quality simultaneously.

When the researchers controlled for quality by swapping the same page to a higher position within its matched pair, the probability of citation rose by only 7.9 percentage points — a result that was not considered statistically significant after accounting for multiple tests. A separate test set of 56 pairs, where only order was switched, produced an estimate of 0.0 points, with a 95% confidence interval from -5.4 to +5.4.

Structured Formatting Showed a Modest, Inconclusive Lift

Pages reformatted with headings and lists received an average of 0.50 more citation markers per answer than their plain-paragraph counterparts, with a 95% confidence interval from 0.20 to 0.84. The total citation count per answer did not increase, suggesting the effect represents a redistribution of credit toward the reformatted page rather than a net gain.

The primary pre-registered outcome — whether a page was cited at all — showed structured text increased citation likelihood by 4.5 percentage points, with a 95% interval from -1.4 to +10.4. The paper notes this result is not conclusive and that the study was only reliably powered to detect effects of approximately 8.5 percentage points or more. The authors framed their findings as an attribution warning rather than a content tactic: "This is an attribution-sensitivity warning, not an optimization tactic."

Model Randomness Is a Significant Confounder

When 120 responses were rerun with identical inputs, citation decisions for the target page changed in 15% of cases — roughly one in seven. The researchers estimate that approximately 45% of the variation in a single run's effect is attributable to model randomness, leading them to recommend that citation tests be run multiple times with consistency metrics reported alongside results.

This aligns with earlier findings: SparkToro reported in January that ChatGPT and Google's AI Overviews each produced the same brand list less than 1% of the time when given the same prompt repeatedly.

Broader Implications for AI Search Strategy

The study raises methodological questions about how practitioners interpret correlational vendor reports. An Ahrefs report from May found that pages cited by AI were about three times more likely to include JSON-LD schema — but, as the paper implicitly highlights, correlation does not confirm that adding schema increases citations. The study also cannot confirm whether reformatting a live page boosts citations, since the rewrites only applied to text already retrieved, leaving crawling, retrieval, and ranking processes outside the experiment's scope.

The authors plan further research across different search providers and models, recommending that future studies test each scenario multiple times and measure both citation frequency and whether a page is cited at all. A rerun of the saved searches on Grok 4.3 found that structured rewrites leaned in the same direction, though fewer than half of Grok's initial replies followed the correct citation format.

The paper has not been peer-reviewed. The study covers one GPT-5.4 search agent using Exa as its search provider, based on offline replayed conversations.

Prepared with AI assistance and reviewed by the editorial team.

Sources

Continue reading

More in AI Discovery →
AI Discovery

NYT Filings Reveal OpenAI Knew of Paywall Workaround; Microsoft Executive Called AI Training 'Astonishing Theft'

Newly unsealed court filings in The New York Times' copyright lawsuit against OpenAI and Microsoft allege that senior figures at both companies were aware of efforts to access paywalled content and use copyrighted material to train AI models. Internal documents cite a Microsoft executive describing the practice as an "astonishing theft of unprecedented proportions" and internal data suggesting AI-powered search could reduce publisher traffic by up to 93%.

3 min read
AI Discovery

AI Search Is Reshaping Retail: What Brands Must Do Now

Conversational AI tools are fragmenting the traditional retail search journey, forcing brands to rethink how they structure product data, create content, and measure visibility. Experts say success now hinges on verified product knowledge graphs, schema markup, and cross-functional teams — not keyword density and backlinks.

7 min read