Method5 min read

You asked once and the model named you. The honest reading is 21% to 100%.

Answer engines give a different answer to the same question depending on how busy the server was. That makes every visibility number a proportion estimated from a sample, and sample size decides whether you have measured anything at all.

One check that named you is consistent with a true rate of 20.7%

The most common way a brand finds out where it stands in AI search is that somebody asks the question once, reads the answer, and either relaxes or panics. That single answer then becomes the number everyone repeats in the next meeting.

Treat it as what it is. One question, asked once, with a yes or no outcome, is a sample of size one from a coin whose bias you do not know. The Wilson score interval is the standard way to put bounds on that kind of proportion. Run it on one success out of one trial at 95% confidence and the answer is 20.7% to 100%.

In plain terms: a model that names you in a single check is entirely consistent with a model that would name you in one answer out of five. The screenshot cannot tell those apart. Neither can you.

The miss is worse. Zero out of one gives an interval of 0% to 79.3%, and teams routinely rewrite a content plan off exactly that.

The same prompt returns a different answer because of who else was asking

This would not matter if engines were deterministic. They are not, and the reason is more specific than the usual hand wave about randomness. OpenAI states it plainly in its own API guidance: Chat Completions are non-deterministic by default, meaning model outputs may differ from request to request. A seed buys you what the same page calls mostly deterministic output, and OpenAI exposes a system fingerprint so you can tell when a backend change has moved the ground under you.

Thinking Machines Lab measured it. In September 2025 the lab sampled 1,000 completions from Qwen3-235B-A22B-Instruct-2507 at temperature zero, which is greedy sampling and should in theory be perfectly repeatable, using the prompt "Tell me about Richard Feynman". They got 80 unique completions, the most common appearing 78 times. All 1,000 runs agreed for the first 102 tokens and split at the 103rd: 992 continued with "Queens, New York" and 8 with "New York City".

The cause they identify is not GPU randomness. The kernels are run to run deterministic, but they are not batch invariant, so a request gets slightly different arithmetic depending on the batch it lands in, and the batch size is set by server load. Their conclusion is worth reading twice: the primary reason nearly all LLM inference endpoints are non-deterministic is that the load, and therefore the batch size, varies non-deterministically.

Which means the answer you screenshotted partly depends on how many strangers were querying the same endpoint at the same second. That is not a signal about your brand.

Twenty samples still leaves a forty point range

Once you accept that a visibility figure is a proportion estimated from a sample, the next question is how big the sample has to be. The Wilson interval answers it, and the answer is humbling. All of these assume the model names you 60% of the time:

  • 5 samples, 3 mentions: 23.1% to 88.2%. A 65 point range. You know the brand exists.
  • 20 samples, 12 mentions: 38.7% to 78.1%. A 39 point range, which is most of the scale.
  • 40 samples, 24 mentions: 44.6% to 73.7%. Now you can rule out being invisible.
  • 100 samples, 60 mentions: 50.2% to 69.1%. Still 19 points wide.
  • 366 samples: the first point where the interval tightens to plus or minus 5 points. Getting to plus or minus 3 takes 1,021.

Most reported movement in an AI visibility dashboard has not happened

Here is where the arithmetic starts costing money. Sample one prompt 40 times in week one and see 24 mentions, then 40 times in week two and see 32. That reads as a jump from 60% to 80%, and it will be on a slide by Friday.

The two intervals are 44.6% to 73.7% and 65.2% to 89.5%. They overlap. A 20 point swing on 40 samples is not yet a result, and the correct thing to tell the room is that nothing has been established. Do that a few times and people stop chasing the noise.

This is why we gate alerting on the interval rather than the point estimate in Proofsource, and why every visibility percentage in our reports carries a Wilson interval beside it. An alert that fires on a number crossing a line fires constantly on a system this noisy. One that fires when the intervals separate fires when something happened.

The same logic applies to zero. No mentions in 20 samples leaves an upper bound of 16.1%, and no mentions in 40 leaves 8.8%. If you want to claim an engine never names you, 40 samples is roughly where that claim starts to mean something.

Use the Wilson interval, not the one you were taught

One technical note, because the obvious alternative fails in exactly the cases that matter here. The interval most people learned is the Wald interval, built from the plain normal approximation. On one success out of one trial it reports 100% to 100%, and on zero out of one it reports 0% to 0%. It tells you a single check is certain, which is the precise error we are trying to fix.

This is well documented. Brown, Cai and DasGupta's review in Statistical Science found that the erratic coverage of the Wald interval is far more persistent than is appreciated, that common textbook prescriptions about when it is safe are misleading and defective and cannot be trusted, and they recommend Wilson for small samples. Answer engine sampling is small samples, because every prompt you track has its own count.

What to do on Monday

None of this requires buying anything. It requires changing what you write down.

  • Stop reporting a single check. If a screenshot is the evidence, the finding is that the brand can appear, and nothing more.
  • Pick a sample size before you look at the data, and keep it the same every period. Changing n between readings makes two numbers incomparable even when both are honest.
  • Report the interval, not the point. A row reading 60% with a range of 45 to 74 is harder to over-read than a row reading 60%.
  • Set the alert threshold on interval separation, and spread the samples across days. Two overlapping intervals are not a change, however different the midpoints look.

Where this could be wrong

The Wilson figures above are exact arithmetic on stated inputs, so the numbers are not in doubt. The modelling assumption behind them is. Wilson intervals assume independent trials with a fixed success probability, and repeated queries to a live engine are not cleanly independent: personalisation, caching, session context and a model version shipping mid-week all violate it. Treat these intervals as a floor on your uncertainty, not a full accounting of it.

The Thinking Machines experiment ran a self hosted open weights model on the lab's own stack, not ChatGPT or Gemini or Perplexity. It establishes the mechanism, not the variance of any commercial product, and none of those vendors publish theirs.

The 60% figure in the sample size table is illustrative. We chose it because intervals are widest near the middle, which makes it the conservative case. A brand sitting at 5% or 95% needs fewer samples for the same absolute precision.

What we would defend is the direction of the error. Small samples, single checks and point estimates without intervals all push the same way, toward believing you have measured something when you have not.

Sources

Where the outside numbers come from

  1. Defeating Nondeterminism in LLM Inference · Thinking Machines Lab
  2. Advanced usage: reproducible outputs · OpenAI
  3. Interval Estimation for a Binomial Proportion · Statistical Science, Institute of Mathematical Statistics
  4. Find information in faster and easier ways with AI Overviews in Google Search · Google Search Help
All notes

Questions

How many times should I ask each prompt?

It depends on the decision. If you only need to know whether an engine ever names you, 20 to 40 samples is enough to make a zero meaningful. If you want to detect a 10 point change between periods, you are into the hundreds per prompt. Decide the question first, then size the sample, and never let the sample size drift between readings.

Does setting temperature to zero fix this?

No. Thinking Machines Lab sampled 1,000 completions at temperature zero and still got 80 distinct outputs, because the non-determinism comes from batch size varying with server load rather than from the sampler. Consumer surfaces like AI Overviews and ChatGPT do not expose temperature to you anyway.

Do the search engines themselves admit to this?

Google's own help page for AI Overviews tells ordinary searchers to ask multiple versions of your question to get the best answers, and says flatly that AI Overviews can and will make mistakes. Consumers are being told to ask more than once. Brands are still checking once.

Why the Wilson interval and not the standard one?

The Wald interval collapses to zero width at 0% and 100%, which is common when you sample a single prompt a handful of times, and its coverage is unreliable well beyond that. Brown, Cai and DasGupta recommend Wilson for small samples, and answer engine sampling is small samples by construction.

Keep reading

Related notes

Your own numbers beat our best post.

The shortlist in your category is being written right now, whether anyone reads this page or not. A free trial, 25 questions on four engines today and tomorrow, tells you if your name is in it.