Method4 min read

Your error bars overlap and the 12 point drop is real

Every dashboard shows a percentage with a bar around it, and everyone reads two of those bars overlapping as nothing happened. That check is far more conservative than anybody intends.

Nobody tells you what to do with the two numbers already on your screen

Your AI visibility number moved. Last week's run said 42% of your prompts named you, this week's says 30%. Twelve points. Before you go hunting for what broke, there's a smaller question in the way: did anything actually happen?

The published advice is thin, and all of it is about your next measurement. Run more samples. Track weekly instead of daily. Keep the prompt set fixed. Sound enough, and no help with the two readings you already have. So most people fall back on what the dashboard makes easy and check whether the error bars overlap. That's the wrong test, and it's wrong in the direction almost nobody expects.

Two 95% intervals overlap more than 99% of the time when nothing has changed

Mark Payton, Matthew Greenstone and Nathaniel Schenker worked this out in the Journal of Insect Science in 2003. Take two samples from populations that are genuinely identical and put a 95% confidence interval on each. How often do those intervals overlap? Their large-sample calculation puts it above 99%, and their simulation, 10,000 pairs at n=10, came out at 0.995.

So checking whether 95% intervals overlap isn't a 5% test. It's closer to a 1% test. You'll almost never call a difference real, including when it is.

The same paper runs the opposite case. Plain standard error bars overlap 84.3% of the time when nothing is going on, so eyeballing those is a test with a false alarm rate near 15%. One time in six you go looking for a cause that isn't there.

They give the correction too. For an overlap check to behave like a 5% test, with standard errors roughly equal, the intervals want to be about 83% or 84% wide. At n=10 their 84% intervals overlapped 94.9% of the time, which is the 5% you were after.

One drop, three verdicts, and the industry uses the quietest one

Put real arithmetic on that 42% to 30% move. Say it's 200 prompts against one engine, 84 naming you in the first run and 60 in the second.

  • 95% Wilson intervals: 42% sits in 35.4% to 48.9%, and 30% sits in 24.1% to 36.7%. They overlap by about a point. Verdict, nothing to see.
  • A standard two-proportion test on the same two numbers: z = 2.50, p = 0.012. Verdict, real.
  • 84% Wilson intervals, per the correction above: 37.2% to 47.0% against 25.7% to 34.7%. Cleanly separated. Verdict, real, and now it agrees with the test, which is the point of the adjustment.

The same prompts run twice is a paired measurement, and pairing is where the power is

There's a bigger mistake under all three. A two-proportion test treats the runs as independent samples. They aren't. Re-run the same 200 prompts against the same engine and every prompt got measured twice, so the comparison belongs prompt by prompt.

Most prompts won't have moved, and the ones that named you both times or neither time tell you nothing about a change. The information is entirely in the prompts that flipped.

Say 30 prompts dropped you and 6 picked you up. Same 24 prompt net, same 12 points. McNemar's test on those 36 flips, the standard test for paired binary data, gives a chi-square of 16.0 on 1 degree of freedom, so p is about 0.00006. The independent test on identical data said 0.012. Pairing threw out the variance that comes from prompts differing from each other, and left the variance you actually asked about.

Change your prompt set between runs and you give that up, and no statistic hands it back. Keep the set fixed and start new prompts as their own series.

Our alerts use the conservative gate too, and here's what that costs you

Our alert rules need two things. The Wilson intervals on the old and new score must not overlap, and the move must clear a threshold you set in a direction you picked. The interval check runs first, so if the bars touch nothing fires, whatever the threshold says.

That's the conservative gate and we chose it on purpose. An alert nobody trusts is worse than no alert. By the numbers above it's roughly a 1% test though, so the 12 point drop in the worked example passes in silence. A real cost, and we'd rather write it down than let you find it.

Underneath the arithmetic there's a floor that matters more. Re-run an identical prompt set against a site nobody touched and the number still moves, because the engines don't repeat themselves. OpenAI's own documentation says it outright: chat completions are non-deterministic by default, and even with a seed set, their system "will make a best effort to sample deterministically" and "determinism is not guaranteed". Across our own sampling we've watched 9 to 27% of the named set turn over from one day to the next, and where it lands depends on the engine. That's with nobody publishing anything.

Which is why Proofsource runs a test-retest report: the same fixed prompts, the same engine, twice a day, on a site that hasn't changed. It gives you the size of your own noise before you set a threshold on top of it. Skip that and your threshold is a guess.

Where this could be wrong

Payton and colleagues derived their result for means from normal populations, assuming equal standard errors and a large sample. Visibility is a proportion, and a Wilson interval goes asymmetric near 0 and 100%. The direction carries over cleanly, 95% overlap is far too conservative for a 5% test, but treat 83 to 84% as a rule of thumb for our case and not a proof about it. Near the extremes, run the test itself.

The worked example is arithmetic, not a measurement. The 200 prompts, the 84 and 60, the 30 flips against 6 are numbers we picked to be self-consistent and to show the gap. Yours will differ, and the split is the thing to look at.

Statistical significance and "you caused this" are separate claims, and only the first one is in scope here. A p of 0.00006 against sampling noise says the engine's behaviour changed between the two runs. It says nothing about why. With churn where we measure it, a significant move can still be the engine drifting on its own.

The 9 to 27% figure is ours. It's what we've seen across the engines we sample, not a published constant, and it's no substitute for measuring your own.

Sources

Where the outside numbers come from

  1. Overlapping confidence intervals or standard error intervals: what do they mean in terms of statistical significance? · Payton, Greenstone and Schenker, Journal of Insect Science 3:34 (2003)
  2. McNemar Test · NIST Information Technology Laboratory, Dataplot reference manual
  3. Advanced usage: reproducible outputs · OpenAI API documentation
  4. Chat API reference, the seed parameter · OpenAI API documentation
All notes

Questions

What should I set my alert threshold to?

Measure your own noise before you pick a number. Run the same prompts against the same engine on a site you haven't changed, and whatever spread that produces is your floor. Put the threshold above it. If you want the alert to catch real moves at the sample sizes most teams run, test the prompts that flipped instead of checking whether two 95% intervals separate.

Should I switch every chart to 84% intervals?

No. The interval on a single score is doing a different job from a comparison between two. 95% is the right width when you're reporting one number and want an honest range around it. The 83 to 84% width is for the narrow case of deciding whether two intervals pulling apart means the difference is real. Show the 95% interval, and run a test for the comparison.

How many prompts do I need before a change is readable?

It depends on the size of the change and how much of your set flips, not on a magic number. Pairing helps a lot. In the example here, 36 flipped prompts out of 200 carried the entire result. If you re-run the same set, you need enough prompts that a genuine move produces more flips than your ordinary day-to-day churn, which is the argument for measuring that churn first.

Does any of this apply to share of voice against a competitor?

The arithmetic applies to any percentage built from a count of prompts, share of voice included. The pairing argument applies whenever the two readings come from the same prompt set. It doesn't apply when you compare your score to a competitor's inside one run, because that's a single sample, and the two scores are not independent of each other either.

Your own numbers beat our best post.

The shortlist in your category is being written right now, whether anyone reads this page or not. A free trial, 25 questions on four engines today and tomorrow, tells you if your name is in it.