One check that named you is consistent with a true rate of 20.7%
The most common way a brand finds out where it stands in AI search is that somebody asks the question once, reads the answer, and either relaxes or panics. That single answer then becomes the number everyone repeats in the next meeting.
Treat it as what it is. One question, asked once, with a yes or no outcome, is a sample of size one from a coin whose bias you do not know. The Wilson score interval is the standard way to put bounds on that kind of proportion. Run it on one success out of one trial at 95% confidence and the answer is 20.7% to 100%.
In plain terms: a model that names you in a single check is entirely consistent with a model that would name you in one answer out of five. The screenshot cannot tell those apart. Neither can you.
The miss is worse. Zero out of one gives an interval of 0% to 79.3%, and teams routinely rewrite a content plan off exactly that.
The same prompt returns a different answer because of who else was asking
This would not matter if engines were deterministic. They are not, and the reason is more specific than the usual hand wave about randomness. OpenAI states it plainly in its own API guidance: Chat Completions are non-deterministic by default, meaning model outputs may differ from request to request. A seed buys you what the same page calls mostly deterministic output, and OpenAI exposes a system fingerprint so you can tell when a backend change has moved the ground under you.
Thinking Machines Lab measured it. In September 2025 the lab sampled 1,000 completions from Qwen3-235B-A22B-Instruct-2507 at temperature zero, which is greedy sampling and should in theory be perfectly repeatable, using the prompt "Tell me about Richard Feynman". They got 80 unique completions, the most common appearing 78 times. All 1,000 runs agreed for the first 102 tokens and split at the 103rd: 992 continued with "Queens, New York" and 8 with "New York City".
The cause they identify is not GPU randomness. The kernels are run to run deterministic, but they are not batch invariant, so a request gets slightly different arithmetic depending on the batch it lands in, and the batch size is set by server load. Their conclusion is worth reading twice: the primary reason nearly all LLM inference endpoints are non-deterministic is that the load, and therefore the batch size, varies non-deterministically.
Which means the answer you screenshotted partly depends on how many strangers were querying the same endpoint at the same second. That is not a signal about your brand.
Twenty samples still leaves a forty point range
Once you accept that a visibility figure is a proportion estimated from a sample, the next question is how big the sample has to be. The Wilson interval answers it, and the answer is humbling. All of these assume the model names you 60% of the time:
- 5 samples, 3 mentions: 23.1% to 88.2%. A 65 point range. You know the brand exists.
- 20 samples, 12 mentions: 38.7% to 78.1%. A 39 point range, which is most of the scale.
- 40 samples, 24 mentions: 44.6% to 73.7%. Now you can rule out being invisible.
- 100 samples, 60 mentions: 50.2% to 69.1%. Still 19 points wide.
- 366 samples: the first point where the interval tightens to plus or minus 5 points. Getting to plus or minus 3 takes 1,021.
Most reported movement in an AI visibility dashboard has not happened
Here is where the arithmetic starts costing money. Sample one prompt 40 times in week one and see 24 mentions, then 40 times in week two and see 32. That reads as a jump from 60% to 80%, and it will be on a slide by Friday.
The two intervals are 44.6% to 73.7% and 65.2% to 89.5%. They overlap. A 20 point swing on 40 samples is not yet a result, and the correct thing to tell the room is that nothing has been established. Do that a few times and people stop chasing the noise.
This is why we gate alerting on the interval rather than the point estimate in Proofsource, and why every visibility percentage in our reports carries a Wilson interval beside it. An alert that fires on a number crossing a line fires constantly on a system this noisy. One that fires when the intervals separate fires when something happened.
The same logic applies to zero. No mentions in 20 samples leaves an upper bound of 16.1%, and no mentions in 40 leaves 8.8%. If you want to claim an engine never names you, 40 samples is roughly where that claim starts to mean something.
Use the Wilson interval, not the one you were taught
One technical note, because the obvious alternative fails in exactly the cases that matter here. The interval most people learned is the Wald interval, built from the plain normal approximation. On one success out of one trial it reports 100% to 100%, and on zero out of one it reports 0% to 0%. It tells you a single check is certain, which is the precise error we are trying to fix.
This is well documented. Brown, Cai and DasGupta's review in Statistical Science found that the erratic coverage of the Wald interval is far more persistent than is appreciated, that common textbook prescriptions about when it is safe are misleading and defective and cannot be trusted, and they recommend Wilson for small samples. Answer engine sampling is small samples, because every prompt you track has its own count.
What to do on Monday
None of this requires buying anything. It requires changing what you write down.
- Stop reporting a single check. If a screenshot is the evidence, the finding is that the brand can appear, and nothing more.
- Pick a sample size before you look at the data, and keep it the same every period. Changing n between readings makes two numbers incomparable even when both are honest.
- Report the interval, not the point. A row reading 60% with a range of 45 to 74 is harder to over-read than a row reading 60%.
- Set the alert threshold on interval separation, and spread the samples across days. Two overlapping intervals are not a change, however different the midpoints look.
Where this could be wrong
The Wilson figures above are exact arithmetic on stated inputs, so the numbers are not in doubt. The modelling assumption behind them is. Wilson intervals assume independent trials with a fixed success probability, and repeated queries to a live engine are not cleanly independent: personalisation, caching, session context and a model version shipping mid-week all violate it. Treat these intervals as a floor on your uncertainty, not a full accounting of it.
The Thinking Machines experiment ran a self hosted open weights model on the lab's own stack, not ChatGPT or Gemini or Perplexity. It establishes the mechanism, not the variance of any commercial product, and none of those vendors publish theirs.
The 60% figure in the sample size table is illustrative. We chose it because intervals are widest near the middle, which makes it the conservative case. A brand sitting at 5% or 95% needs fewer samples for the same absolute precision.
What we would defend is the direction of the error. Small samples, single checks and point estimates without intervals all push the same way, toward believing you have measured something when you have not.