Nobody tells you what to do with the two numbers already on your screen
Your AI visibility number moved. Last week's run said 42% of your prompts named you, this week's says 30%. Twelve points. Before you go hunting for what broke, there's a smaller question in the way: did anything actually happen?
The published advice is thin, and all of it is about your next measurement. Run more samples. Track weekly instead of daily. Keep the prompt set fixed. Sound enough, and no help with the two readings you already have. So most people fall back on what the dashboard makes easy and check whether the error bars overlap. That's the wrong test, and it's wrong in the direction almost nobody expects.
Two 95% intervals overlap more than 99% of the time when nothing has changed
Mark Payton, Matthew Greenstone and Nathaniel Schenker worked this out in the Journal of Insect Science in 2003. Take two samples from populations that are genuinely identical and put a 95% confidence interval on each. How often do those intervals overlap? Their large-sample calculation puts it above 99%, and their simulation, 10,000 pairs at n=10, came out at 0.995.
So checking whether 95% intervals overlap isn't a 5% test. It's closer to a 1% test. You'll almost never call a difference real, including when it is.
The same paper runs the opposite case. Plain standard error bars overlap 84.3% of the time when nothing is going on, so eyeballing those is a test with a false alarm rate near 15%. One time in six you go looking for a cause that isn't there.
They give the correction too. For an overlap check to behave like a 5% test, with standard errors roughly equal, the intervals want to be about 83% or 84% wide. At n=10 their 84% intervals overlapped 94.9% of the time, which is the 5% you were after.
One drop, three verdicts, and the industry uses the quietest one
Put real arithmetic on that 42% to 30% move. Say it's 200 prompts against one engine, 84 naming you in the first run and 60 in the second.
- 95% Wilson intervals: 42% sits in 35.4% to 48.9%, and 30% sits in 24.1% to 36.7%. They overlap by about a point. Verdict, nothing to see.
- A standard two-proportion test on the same two numbers: z = 2.50, p = 0.012. Verdict, real.
- 84% Wilson intervals, per the correction above: 37.2% to 47.0% against 25.7% to 34.7%. Cleanly separated. Verdict, real, and now it agrees with the test, which is the point of the adjustment.
The same prompts run twice is a paired measurement, and pairing is where the power is
There's a bigger mistake under all three. A two-proportion test treats the runs as independent samples. They aren't. Re-run the same 200 prompts against the same engine and every prompt got measured twice, so the comparison belongs prompt by prompt.
Most prompts won't have moved, and the ones that named you both times or neither time tell you nothing about a change. The information is entirely in the prompts that flipped.
Say 30 prompts dropped you and 6 picked you up. Same 24 prompt net, same 12 points. McNemar's test on those 36 flips, the standard test for paired binary data, gives a chi-square of 16.0 on 1 degree of freedom, so p is about 0.00006. The independent test on identical data said 0.012. Pairing threw out the variance that comes from prompts differing from each other, and left the variance you actually asked about.
Change your prompt set between runs and you give that up, and no statistic hands it back. Keep the set fixed and start new prompts as their own series.
Our alerts use the conservative gate too, and here's what that costs you
Our alert rules need two things. The Wilson intervals on the old and new score must not overlap, and the move must clear a threshold you set in a direction you picked. The interval check runs first, so if the bars touch nothing fires, whatever the threshold says.
That's the conservative gate and we chose it on purpose. An alert nobody trusts is worse than no alert. By the numbers above it's roughly a 1% test though, so the 12 point drop in the worked example passes in silence. A real cost, and we'd rather write it down than let you find it.
Underneath the arithmetic there's a floor that matters more. Re-run an identical prompt set against a site nobody touched and the number still moves, because the engines don't repeat themselves. OpenAI's own documentation says it outright: chat completions are non-deterministic by default, and even with a seed set, their system "will make a best effort to sample deterministically" and "determinism is not guaranteed". Across our own sampling we've watched 9 to 27% of the named set turn over from one day to the next, and where it lands depends on the engine. That's with nobody publishing anything.
Which is why Proofsource runs a test-retest report: the same fixed prompts, the same engine, twice a day, on a site that hasn't changed. It gives you the size of your own noise before you set a threshold on top of it. Skip that and your threshold is a guess.
Where this could be wrong
Payton and colleagues derived their result for means from normal populations, assuming equal standard errors and a large sample. Visibility is a proportion, and a Wilson interval goes asymmetric near 0 and 100%. The direction carries over cleanly, 95% overlap is far too conservative for a 5% test, but treat 83 to 84% as a rule of thumb for our case and not a proof about it. Near the extremes, run the test itself.
The worked example is arithmetic, not a measurement. The 200 prompts, the 84 and 60, the 30 flips against 6 are numbers we picked to be self-consistent and to show the gap. Yours will differ, and the split is the thing to look at.
Statistical significance and "you caused this" are separate claims, and only the first one is in scope here. A p of 0.00006 against sampling noise says the engine's behaviour changed between the two runs. It says nothing about why. With churn where we measure it, a significant move can still be the engine drifting on its own.
The 9 to 27% figure is ours. It's what we've seen across the engines we sample, not a published constant, and it's no substitute for measuring your own.