Google documents the fan-out in its own words
The mental model most teams still carry is the ranked list. You have a keyword, the engine has an index, and somewhere in between sits a position you can move. Every AI search product shipped in the last two years broke that model, and Google has written down exactly how.
Google's guidance for site owners says AI Overviews and AI Mode may use a query fan-out technique, issuing multiple related searches across subtopics and data sources to develop one response. The I/O announcement from Elizabeth Reid, Google's VP and Head of Search, is more specific about AI Mode: it breaks your question into subtopics and issues a multitude of queries simultaneously on your behalf. Google's stated reason is reach. Fanning out lets Search go deeper into the web than a single traditional query does.
So the thing being answered is not the thing that was typed. By the time a response assembles, the engine has run a set of questions it wrote itself, and your page either turned up in some of them or it did not.
A dozen on an ordinary question, hundreds on a hard one
Google has put a number on it twice. In March 2026 its own interview series asked Dounia Berrada, a Senior Engineering Director on Search, to explain the method. Her summary: AI Mode is basically doing a dozen searches for you in the time it takes to do one. Her example is a photo of a garden, where one upload carries several real questions at once about shade, climate and upkeep, and the model issues a search for each of them.
At the other end of the range, Deep Search in AI Mode uses the same technique taken further. Google says it can issue hundreds of searches, reason across disparate pieces of information, and produce a fully cited report. The fan-out has gone multimodal too. Circle to Search and Lens now break a single image into its separate objects and search them in parallel, so one photo of a room becomes searches for the lamp, the rug and the chair.
A dozen is the number worth keeping, because it describes the ordinary case rather than the showcase.
Every engine decomposes, and none of them agree on how far
This is not a Google-only habit, and the budgets are nowhere near equal. Perplexity publishes the default instructions its Agent API assistant runs with, and they say to decompose complex user queries into discrete tool calls for accuracy and parallelization. The same published default then puts a ceiling on it: make at most three tool calls before concluding.
Three and a dozen and hundreds are different machines. Google's documentation adds that AI Overviews and AI Mode may use different models and techniques, so the answers and the links they show will vary between two surfaces inside the same company.
That variation is the practical problem. A brand can hold the shortlist on an engine that fans out twelve ways and miss it on an engine that fans out three ways, for the same question, on the same day, with the same site and no change to anything.
The subquery you lose on is never the one you tracked
Here is the failure that costs the most. A team picks the head term that matches its product page and watches that. The brand shows up. The dashboard is green.
Meanwhile the fan-out wrote eleven other questions. Price against two named competitors. Whether it works for a team of five. What people who left it moved to. Integration with a tool nobody in marketing thinks about. Each of those has its own results, its own cited sources and its own winners, and they get assembled into the answer a buyer actually reads.
Nothing reports any of it to you. Google counts AI Overviews and AI Mode inside ordinary Web search traffic in Search Console, so the subqueries never appear as their own rows. The only way to see them is to ask the engines the questions a buyer would ask, repeatedly, and record what comes back.
That is the measurement we built Proofsource to run: sample each engine across a prompt set, map the fan-out, track which sources get cited alongside which, and put a Wilson interval on every number so a move has to clear the noise before it counts as a move.
Four ways to measure a question you never asked
None of this needs our product. It needs you to stop treating one prompt as one prompt.
- Write the fan-out yourself before you measure anything. Take your head term and list the ten or twelve questions a buyer has to get answered underneath it: price, alternatives, who it suits, what it connects to, why people leave. That list is your real prompt set.
- Ask each engine separately and keep the answers. AI Mode, AI Overviews, ChatGPT, Perplexity and Gemini will not agree with each other, and the disagreement is the finding rather than a fault in your method.
- Sample repeatedly before you conclude anything. One run tells you what happened once. We see 9% to 27% day-over-day churn in which names appear, depending on the engine, in the categories we sample. A single reading off that base is not a measurement.
- Read the cited sources, not only the mentions. In the category we sampled, listicles were 44% of what ChatGPT cited for a best-of question. If third-party roundups are answering your subqueries, your own site being perfect does not settle it.
Where this could be wrong
The mechanism above is Google and Perplexity describing their own systems, and neither company publishes the actual subqueries behind a given response. A dozen is an engineering director's round summary in an interview, not a measured distribution, and it will move by question, by surface and by month. Treat it as an order of magnitude and nothing finer.
The cap of three tool calls is the published default in Perplexity's Agent API documentation. It describes that documented default and says nothing about what the consumer product does on any particular question.
Our own figures are ours. The 9% to 27% churn and the 44% listicle share are single measurements on the samples we took, in the categories we took them in, and we would not defend either as a general law. Google's own site-owner page was last updated in December 2025, and these products change faster than their documentation does.
What we would defend is the shape of the mistake. Tracking one string against a system that writes its own questions measures the single input you control and none of the ones that decide the answer.