How ChatGPT Chooses Its Sources

How ChatGPT decides which sources to cite: live web search, training data, and the sites it leans on most, like Wikipedia and Reddit. Plain-English guide.
V
Written by Victor
Updated 3 months ago

ChatGPT builds answers from two places: what it learned during training, and live web results it pulls in when it searches. When it cites sources, it leans heavily on a few trusted sites, especially Wikipedia and Reddit. The way to get cited is to be present and credible on the sources it reads, not to talk to the model directly. Here is how it works and what it means for your brand.

Two ways ChatGPT knows things

ChatGPT has two sources of knowledge, and they work differently.

The first is its training data: the large body of text it learned from before it answered any question. You cannot edit this directly, and it updates only when the model is retrained.

The second is live web search. For many questions, ChatGPT runs a real-time search, pulls in current pages, and writes its answer with links to the sources it used (OpenAI, Introducing ChatGPT search). This is where most of the day-to-day citation activity happens, and it is the part you can influence by improving what is on those pages.

One technical detail matters here: ChatGPT can only cite what it is allowed to read. OpenAI's OAI-SearchBot crawler decides which sites can appear in ChatGPT search answers, and sites that block it in robots.txt will not be shown (OpenAI crawler docs). Training uses a separate crawler, GPTBot, which a site can allow or block independently.

The sites it leans on most

ChatGPT does not treat all sources equally. It pulls again and again from a small set of trusted sites.

A Semrush study of more than 100 million citations found that ChatGPT leaned heavily on Wikipedia, Reddit, and editorial sites like Forbes (Semrush, most-cited domains study). Community discussion and reference content carry a lot of weight, because they read as real people comparing options and as established fact.

It also cites a lot, but selectively. Muck Rack's May 2026 study of 25 million AI-cited links found ChatGPT included sources in 96% of responses but averaged only five citations per answer, with Wikipedia its single most-cited domain (Muck Rack, What Is AI Reading?). Five slots per answer means the competition for each one is real.

Company-owned content matters too, but less. A BuzzStream analysis reported by Search Engine Journal found that owned newsroom and editorial content on company domains made up about 18% of ChatGPT's citations, while syndicated press releases over the wire earned almost nothing (Search Engine Journal).

Its sources shift, sometimes sharply

What ChatGPT cites is not fixed. The same Semrush study caught a dramatic swing: ChatGPT's use of Reddit fell from roughly 60% of responses to around 10% within weeks in late 2025, and its use of Wikipedia dropped from about 55% to under 20% over the same period. Semrush read this as the platform reducing over-reliance on any single source.

The lesson is not to chase one platform. It is to build genuine presence across several of the sources ChatGPT reads, so a change to any one of them does not erase your visibility.

What this means for getting cited

You cannot tell ChatGPT to recommend you. It builds answers from evidence it can find and check, so the work is to improve that evidence.

In practice that means being genuinely present where ChatGPT looks: helpful, accurate content on community sites, credible coverage on real editorial outlets, and clear information about your brand on the open web. Quality is the filter. Useful content survives the platform's adjustments; spammy content is what gets discounted.

Where to go next

For the full picture of how AI search works, read what is AEO, GEO and AI search. To see how Cited builds that presence for you, see how Cited works.

Did this answer your question?