Playbook4 min read

ChatGPT named the brand 78% of the time it searched, and 6% when it didn't

Before you rewrite a page, find out whether the model read any page at all. That one fact moved the mention rate further than anything else we measured.

Two thirds of our ChatGPT answers never searched the web

Search for why ChatGPT won't mention your brand and you get a consistent list: build authority, get into the roundups it trusts. Every item assumes the model read something about your category and ranked you below somebody else. A lot of the time it read nothing.

Between 22 August and 1 October 2026 we scored 364 completed answers where we could check both halves: which brand the answer named, and which pages it cited. Four brands across three accounts in adjacent B2B software categories, on the two engines that return citations. Google AI Overviews, and our ChatGPT engine, which calls OpenAI's API with the web search tool attached. It's a narrow slice, so read every number here as ours.

  • 226 of those answers came from the ChatGPT engine, and 215 of them stored a count of how many web searches the model ran.
  • 147 of the 215 ran zero searches, so nothing was retrieved at all. That's 68%.
  • 68 ran one or two searches.

Searching moved the mention rate from 6% to 78%

Of the 147 answers that searched nothing, 9 named the brand. That's 6.1%, with a 95% confidence interval of 3.3% to 11.2%. Of the 68 that searched, 53 named it: 77.9%, interval 66.7% to 86.2%. The intervals don't come close to touching.

The obvious objection is that these are different questions. A model is likelier to search a commercial comparison than a definitional one, and likelier to name vendors in the first kind. So we took the 9 prompts that had been asked repeatedly and had gone both ways, and compared the same question against itself. Browsing runs named the brand in 18 of 23 answers, 78.3%. Non-browsing runs, 7 of 29, 24.1%. Smaller gap, same direction, intervals still clear of each other.

Treat 24% instead of 6% as the honest floor. Whether the model looked anything up is the biggest thing we found separating an answer that names you from one that doesn't, and it has nothing to do with your page.

The search is the model's call, taken one prompt at a time

The advice circulating about this is a few years out of date. One widely read guide tells you to run each prompt twice, once in default chat and once with browsing on, as though the reader picks. Another explains that ChatGPT "doesn't retrieve specific documents" and works from training patterns, which describes only the answers where no search happened.

OpenAI's API documentation is plainer. Attaching the web search tool makes it available, and in their words the model can choose to search the web or not based on the content of the input prompt. With tool_choice set to auto, search is optional; you need tool_choice required if a search must run. Nobody flips a switch. The model decides, prompt by prompt, and it can decide differently on the same question twice.

Which means your visibility figure is two populations added together. Some answers looked and left you out. Some never looked. Different problems, different fixes, and one percentage hides both.

Google AI Overviews has no version of not looking

The other engine behaved nothing like it. Of the 138 AI Overview answers we scored, 81 didn't name the brand, and all 81 of them cited other pages anyway. A median of 8 sources per answer, up to 16.

That makes a miss on AI Overviews easy to read. Eight pages were pulled for that question and none was yours. Google's guidance says a page is eligible as a supporting link if it's indexed and can be shown with a snippet, with no extra requirements, and that meeting every requirement still doesn't mean Google will crawl, index or serve it. Google also says AI Overviews appear only when its systems judge them additive to classic Search, so they often don't trigger. That non-appearance is a third outcome, and it isn't in these counts.

94 of the 97 answers that cited the brand's own page named it too

Read from the retrieval side, the relationship is lopsided. Across both engines, 97 answers put a page from the brand's own domain in the source list, and 94 of them named the brand: 96.9%, interval 91.3% to 98.9%. On the ChatGPT engine it was 44 out of 44. Among browsing answers that cited no page of the brand's, the rate fell to 9 of 24.

It doesn't run the other way. 119 answers named the brand, and only 94 of those had a page of its own in the sources, 79.0%. One in five named the brand with nothing of the brand's retrieved, which is the model working from what it already knows.

Getting one of your own URLs into the retrieval set was close to sufficient for a mention in our data, and never necessary. From 97 answers we can't call it causal.

What we'd actually do

In order, cheapest first.

  • Split your mention rate by whether the answer retrieved anything before you act on it. One number over both populations is the mistake.
  • For the answers that retrieved nothing, rewriting a page this week won't move them. That's the model's prior, and it shifts on the timescale of the wider web talking about you.
  • For the answers that did retrieve, read the sources they used. On AI Overviews that's about eight pages per question, and it names exactly who got pulled instead of you.
  • Ask any tool quoting you an AI visibility score which of these two things it's counting. If it can't answer, the number is an average of two different failures.

Where this could be wrong

The biggest caveat is ours, not the engine's. Our ChatGPT engine attaches the web search tool and leaves tool_choice on auto, forcing a search only on prompts that read as commercial or recency-driven, so the 68% that searched nothing is partly a product of our own policy. A real ChatGPT session routes its own searches and may browse more often than our sampler provokes. The within-prompt comparison is unaffected, since both sides ran under the same policy, and that's the number we'd stand behind.

Four brands and three accounts is thin, and two of those brands supplied most of the answers. The scored set is a subset too: it covers runs where mention extraction had produced a row for the brand, 226 of 1,047 ChatGPT runs. The search counts come from search-call items in the stored response, including any that failed, so they're an upper bound.

One more, which OpenAI documents itself: the inline citations aren't the full retrieval. Their sources field returns every URL the model consulted, and they note it's often longer than the citation list. We read citations, so "your page wasn't in the sources" means "not in the sources we can see", and the real retrieval rate for an own domain sits above what we measured. The AI Overview payloads reach us through DataForSEO and we didn't re-render those searches ourselves.

Sources

Where the outside numbers come from

  1. Web search (Responses API guide) · OpenAI API documentation
  2. AI features and your website · Google Search Central
All notes

Questions

How do I tell whether ChatGPT searched before it answered?

In the product, an answer that searched carries sources; one that didn't carries none. Through the API, the response contains a web search call item when a search ran, and the count of those items is the clean signal. In the ChatGPT app, visible inline citations are the tell. No citations usually means no retrieval, so the answer came from the model's existing knowledge.

My AI visibility score dropped. Does this explain it?

It might, and that's the problem. If the share of answers that searched changed between two runs, your score moves without anything about your site changing. We'd check the retrieval mix before reading a drop as a content or authority problem, and we'd want a confidence interval on both numbers before calling it a change at all.

If the model answers from memory, is there anything I can do?

Not quickly. Those answers reflect what the wider web said about your category over a long period, so they respond to sustained third-party coverage and not to a page edit. The useful move is to stop spending content effort against them and put it where retrieval is actually happening, which in our sample was every single Google AI Overview answer.

Does Proofsource separate these two cases?

It stores what it needs to. Every sampled answer keeps the engine's full normalised response, including the citations it emitted and the number of web searches it ran, so both populations are distinguishable in the data we hold. The dashboard reports citations and visibility; splitting the headline mention rate by retrieval is not a view it ships today.

Keep reading

Related notes

Your own numbers beat our best post.

The shortlist in your category is being written right now, whether anyone reads this page or not. A free trial, 25 questions on four engines today and tomorrow, tells you if your name is in it.