AI Search Tools Draw on Dozens of Data Sources, From Google to Yelp, Analysis Finds
A tiered reference table published by SEO consultant Chris Green maps the data sources currently used by AI search tools such as Google Gemini, Microsoft Copilot, and ChatGPT, ranging from confirmed real-time grounding feeds to historical pretraining corpora, and urges search practitioners to broaden their focus beyond traditional web rankings.
AI-powered search tools are drawing on a far wider range of data sources than traditional search engines, according to a reference guide published on 22 September 2026 by Chris Green, Technical Director and Senior Consultant at Torque Partnership, and syndicated by Search Engine Journal.
Green compiled a tiered table categorising sources by their confirmed or inferred role in AI systems, covering web discovery, shopping feeds, local data, knowledge bases, community content, news, developer resources, and travel commerce. The table distinguishes between sources confirmed for real-time retrieval and grounding (Tier 1), confirmed for training or licensing (Tiers 2–3), and those with strong but unconfirmed evidence (Tier 4).
Real-Time Grounding Sources
Among the Tier 1 confirmed sources, Green notes that Google's Gemini API uses live web content for grounding, with Google Search connecting the model to real-time results and inline citations. Google Maps is listed alongside Search as a documented grounding tool giving models geospatial context. For local commerce, Yelp is confirmed to license its reviews, photos, and business information to OpenAI for real-time local recommendations, with functionality extending to in-chat actions such as table bookings and quote requests; Green cites Yelp's 10-Q as confirming the partnership is live.
In the travel sector, Green reports that Google's Hotel Center feeds underpin hotel options shown in Gemini and AI Mode, including real-time pricing, and notes that in August 2026 Google added hotel booking inside AI Mode completed with Google Pay, making it a grounding-plus-actions use case.
Wikipedia and Wikimedia are listed as Tier 1 sources used both in pretraining and as live reference corpora. Reddit is assigned to both Tier 1 and Tier 2: the Google–Reddit deal is described as giving access to the Reddit Data API for real-time structured content and allowing Reddit content to be displayed across Google products, while also permitting Google to use Reddit posts to train its AI models, reportedly at around $60 million per year. Green flags that Reddit is reportedly weighing whether to renew the agreement and advises treating it as an unstable source.
Licensed and Training Sources
On the publisher licensing side, Green notes that OpenAI holds multiple explicit licensing partnerships including with the Financial Times, Axel Springer, the Associated Press, and News Corp, with terms differing per partner on training, grounding, and attribution. The Axel Springer deal is noted to include otherwise paywalled material in answers.
For shopping, OpenAI's agentic commerce specification allows merchants to share a regularly refreshed CSV or JSON feed of identifiers, descriptions, pricing, inventory, media, and fulfilment information so ChatGPT can surface products accurately, with refreshes accepted as often as every 15 minutes. Shopify catalogue data is cited as already integrated into ChatGPT.
Developer sources including GitHub's public BigQuery dataset — restricted to Apache, BSD, and MIT-licensed projects — and Stack Overflow are listed as confirmed training or licensing sources at Tier 2.
Historical Pretraining Corpora
Green identifies Common Crawl as a major historical pretraining source, noting that GPT-3 used filtered Common Crawl as roughly 60 percent of its sampling mixture and LLaMA 1 reported 67 percent. The C4 dataset, a cleaned derivative of Common Crawl, accounted for 15 percent of LLaMA's pretraining mixture according to the analysis.
Uncertainty and Practitioner Implications
Green acknowledges the inherent instability of the landscape, writing that the table may date badly, describing it as "an occupational hazard of 'AI Search' at the moment." Several entries — including npm and PyPI package registries — are flagged as weakly evidenced, with Green noting no disclosed agreement or documented retrieval use was found for those sources.
The analysis recommends that practitioners study AI-generated results for queries their customers are likely to use in order to identify gaps, and cautions that the relevance of specific sources varies by market — citing Yelp's limited reach in the United Kingdom as an example where other regional data providers may carry greater weight.
Prepared with AI assistance and reviewed by the editorial team.