Every guide says start a fresh chat, then describes a five-turn buyer
The standard way to check whether an answer engine names your brand is to write a set of buyer questions, run each in a brand new chat with no prior context, and record who got named. We do this too, for a sound reason: a chat you have already steered toward your own brand tells you what a warmed-up model says, not what a cold buyer sees.
Then read what the same guides say the buyer is doing. Someone choosing a vendor asks what to look for, mentions a constraint, asks for a comparison, asks where to buy. Google treats that as the point of the product: its help page for AI Mode says you can ask anything and get a response, with the ability to go deeper through follow-up questions. The interaction these engines are built around runs over several turns. Every number in every AI visibility report is the number at turn one.
A 200,000 conversation experiment found the drop arrives at turn two
One large controlled measurement of this exists. Philippe Laban and Jennifer Neville at Microsoft Research, with Hiroaki Hayashi and Yingbo Zhou at Salesforce Research, published it on 9 May 2025. They took 600 fully specified instructions across six generation tasks, split each into pieces, and had a simulated user reveal at most one piece per turn. Fifteen models, over 200,000 simulated conversations.
Models scoring around 90% when handed the whole instruction in one message scored around 65% when the same information arrived across turns. Averaged over models and tasks the degradation was 39%, and every model degraded on every task.
Splitting an instruction could simply lose information in the rewrite. They controlled for it: concatenating the same pieces back into one bullet list scored 95.1% of the original. The rewriting costs nothing, the turns cost everything. Varying how finely they split, from two pieces up to eight, the damage appeared at two and stayed flat after that.
The models did not get less capable. They got less repeatable.
This is the part that matters if you report a percentage to anyone. The paper splits performance into aptitude, the 90th percentile score across repeated runs, and unreliability, the gap between the 90th and the 10th. Multi-turn aptitude fell about 16%. Unreliability rose 112%. The best multi-turn run is nearly as good as the best single-turn run and the worst is far worse, with the spread on a fixed instruction averaging around 50 points.
Neither usual remedy helped. o3 and DeepSeek R1 got lost like everything else, and the authors blame length: reasoning models wrote answers about 33% longer, carrying more of their own assumptions forward. Lowering the temperature cut unreliability by 50 to 80% in the single-turn settings and did nothing useful in the multi-turn ones, where it sat near 30% even pinned to zero.
Read that against a visibility report. A shortlist read once at turn one is already one sample of a noisy process, which is why we publish an interval beside every percentage. A shortlist read once at turn three is one sample of a process with twice the spread.
Your turn-one shortlist becomes an input to turn two
OpenAI documents the mechanism. Each request is independent and stateless, and a multi-turn conversation works by sending the earlier messages back in alongside the new one. Even when you chain responses by ID, OpenAI notes, every previous input token is billed again as an input token. The history is not remembered, it is resent, and that includes the assistant's own earlier answer.
So when a buyer asks a follow-up, the model reads its own turn-one shortlist as part of the input, and Laban and colleagues name over-reliance on previous answer attempts as one of four causes of the effect. Which points at something we cannot yet prove: at turn one you compete for a slot, at turn two you argue with a list the model has already written. Nobody has measured that on brand shortlists. It is the reading we would bet on.
What we measure, and the number none of us has
Proofsource samples single turn by construction. We ask each prompt repeatedly, score visibility with a Wilson interval rather than a point, and run a test-retest report that asks the same question two days running. None of that samples a follow-up as a follow-up. No tool we know of does, ours included, and we would rather write that down than imply coverage we lack.
What we can show is the neighbouring effect: change the question and the named set changes. In one full report on a single brand in a business software category, the category question returned fourteen vendors and the brand was not among them. Asked about by name, the same engine called it one of the leading platforms in its category and rated it four to five stars. Across sixteen phrasings of that one commercial question, six competitors appeared in eleven or more, and the brand in at most two. That is ours, one brand and one engine, and it is a phrasing result rather than a turn result.
- Write the follow-up as its own prompt. "Best project management tool" and "best project management tool for a twelve person team on no budget" are two prompts, not one.
- Put the constraint in the first message. That is the paper's advice to users, and the only way to get a reading you can compare next month.
- Stop treating a screenshot of a three-turn conversation as evidence, either way. It is one run of the least repeatable setting anybody has measured.
Where this could be wrong
The largest gap is the transfer. The six tasks are Python functions, SQL queries, API calls, grade-school arithmetic, table captions and query-focused summaries. None of them is recommend a vendor. The paper establishes that models lose track of requirements arriving over turns and get far less repeatable when they do. It does not establish that brand shortlists behave the same way. We think they do, and we have not measured it.
Its users are simulated by a small model rather than people, and the tasks are English, text only, and analytical. The authors list all of that, and argue their conditions are benign enough that real degradation is probably larger, which supports the direction and says nothing dependable about the size. These are also 2025 models, and none of the four engines we track publishes a multi-turn reliability figure of its own.
Our sixteen-phrasing numbers are one brand, one category, one engine, one report. They show the question deciding the list, and they are not a turn measurement. We also have no figure for how often a real buying conversation runs past one turn: Google calling follow-ups the point of AI Mode is a design statement, not a frequency.
What we would defend is narrow and a little awkward for us. What every product in this category reports is turn one, what everyone talks about is the conversation, and right now none of us is measuring the difference.