Explainer

Where Does ChatGPT Get Its Information?

ChatGPT's answers come from two places with almost nothing in common: a frozen training corpus and a live web search. Which one answered you decides what, if anything, you can do about it.

Small translucent page cards flowing along a curved ribbon through a glossy glass sphere and emerging as one frosted answer panel

Ask where ChatGPT gets its information and you have really asked two questions, because there are two supplies, so different they might as well belong to separate products — a frozen archive on one side, a live wire on the other. Which of the two produced the answer in front of you changes what that answer is, how current it is, and whether anything can be done about it.

Most of the confusion about AI answers — and most of the bad advice sold about them — comes from not knowing which of the two you are dealing with on a given question.

The frozen half: what the model was trained on

The base model behind ChatGPT was trained on an enormous snapshot of public text: websites, books, encyclopedias, forums, news, documentation. Training ended at a cutoff date, and nothing after that date exists in the model's weights. In this half of the system there is no lookup, no database, no retrieval. The model does not find a fact the way a search engine finds a page; it reconstructs what the public record tended to say, because saying it is what the training process rewarded.

Two properties of this half matter for anyone who cares about being mentioned.

First, it is statistical, which means it favors repetition. A brand described the same way in a thousand places casts a sharper shadow in the weights than a brand described brilliantly in one place. When ChatGPT answers "what are the well-known accounting tools?" from memory, it is, loosely, reciting the consensus of years of public writing about accounting software. If that consensus never included you, the memory answer will not either.

Second, it moves on training timescales. Nothing you publish this month changes a memory answer this month. The earliest it can matter is the next time a model is trained on a snapshot that includes it. This is why the trained half rewards patience and consistency, and cannot be gamed on any shorter timescale.

The live half: when it goes and reads the web

For plenty of questions, though, ChatGPT does not answer from memory. It runs a web search in the middle of composing its reply, reads a handful of results, and builds the answer largely out of what those pages say. You can see this half working: the answer arrives with citations, current prices, dates from this week.

The retrieval plumbing is worth knowing precisely, because it is one of the few parts of this whole system that is documented. OpenAI's crawler for this purpose is called OAI-SearchBot, and OpenAI's own documentation states plainly that sites opted out of it "will not be shown in ChatGPT search answers" — though the same document allows that they may still appear as bare navigational links. At launch OpenAI described ChatGPT search as running on "a mix of search technologies, including Bing"; since then it has been shifting toward its own crawl. Either way, the pool it draws from is the ordinary, indexed public web — the same pages search engines have been ranking all along.

What it does with that pool is selective in a way that should change how you think about content. A search engine shows you ten links and lets you choose. The live half of ChatGPT reads a few pages and composes one paragraph. The pages it happened to read carry absolute weight, and the eleventh-best page in the world carries none. In practice, for commercial questions, the pages it reads are the category's listicles, review sites, head-to-head comparisons, and forums. The answer inherits whatever names those few pages agree on.

Five translucent glass cards stacked vertically, each linked by thin light threads to one glowing answer card above them
A handful of retrieved pages decide what the composed answer says.

How to tell which half answered you

You can usually diagnose an answer at a glance. Citations, links, "according to", prices, and recent dates mean the live half ran. No sources, hedged phrasing, and a certain smooth generality mean you got the archive. The same question can go either way on different days, which is one reason two people comparing their ChatGPT answers so often talk past each other.

This distinction is not academic. A memory answer and a retrieval answer about your brand have different causes, and therefore different remedies.

It searches less often than you think

Here is the part that surprises people who work with the API all day. The consumer product — the free tier your actual buyers use — is configured more conservatively than developer defaults, and it reaches for the web less often. We measure AI answers for a living, and when we pinned our tracking to the consumer tier, the rung a real logged-out user gets, the number of web searches our tracked questions triggered roughly halved compared with developer-default settings.

The decision to search also belongs entirely to the model, and it is not a steady one. Watching Gemini's grounding behavior while building our own tracking pipeline, we saw the same prompt trigger seventy-six searches on one call and four on the next. Assistants are moody about retrieval, and no amount of optimization on your side forces the issue.

The consequence for brands is uncomfortable: a meaningful share of the answers your buyers see are memory answers, untouched by anything you shipped this quarter. Advice that treats every AI answer as a retrieval answer — most "AI SEO" advice, in other words — is optimizing for the half of the system that is easier to influence, and quietly ignoring the other half.

What this means if you own a brand

The two halves respond to different work.

The live half responds to presence in the pages it retrieves. Find out which pages decide the answers in your category — ask the assistant your buyers' questions and read what it cites — and get named there. It also responds to plumbing: an allowed OAI-SearchBot and a clean Bing presence are table stakes, verifiable in an hour.

The frozen half responds to the long record: a public description of your brand that is consistent everywhere it appears, accumulating for years across directories, reviews, press, and your own site. It is slow by construction. That slowness is also a moat — a competitor cannot buy their way into the archive quickly either.

If you want the operational version of this — what to do first, what to skip, how to check whether any of it is working — Bondo has written the playbook, and the fastest way to see your own starting point is to run one free check on a question your buyers actually ask.

FAQ

Does ChatGPT use Google or Bing?

Neither, in the sense people mean. It does not forward your question to a search engine you could use yourself. Its search capability leaned on Bing at launch and has been shifting toward OpenAI's own crawling via OAI-SearchBot; the results it reads come from that pool, not from Google.

How current is ChatGPT's information?

Split the question in two. Memory answers are only as current as the training cutoff, which trails the present by months or more. Retrieval answers are as current as the pages fetched while you waited. If the answer carries citations and today's prices, you are looking at the second kind.

Does ChatGPT read my website directly?

It can, in two ways: your pages may be in the training snapshot, and OAI-SearchBot may fetch them during a live search — if your robots.txt allows it. Neither guarantees you get named. Your site is one voice in the record; for commercial questions, the third-party pages that describe your category usually speak louder.

Why does it give different answers to the same question?

Three reasons stack: the model does not search on every run, the pages retrieved can differ between runs, and generation itself is not deterministic. This is why a single check of "does it mention us" tells you very little, and why anything serious about AI visibility has to measure a rate over repeated runs rather than trust one screenshot.

Track what AI answers,
every day.

Your buyers' questions, asked to every assistant, scored on one rubric.