A mention is close to a fixed property of the brand, the prompt and the engine
Most AI visibility dashboards lead with one number, the share of answers that name you. It gets screenshotted into board decks and it's what people put alerts on. It's also the part of the answer that moves least.
A paper published in June 2026 by Pratyush Kumar at Ranqo puts a size on that. Ranqo sells GEO tooling and is reporting its own production data, which matters for how much weight you give it. The cohort is 102 brands across five engines, March to May 2026, more than 100,000 prompt responses.
For every brand, prompt and platform cell tracked over at least three runs on unbranded category prompts, the mention pattern sorted like this: never mentioned 63.2%, always mentioned 14.3%, mentioned in at least 70% of runs 7.0%, mentioned in under 30% 8.7%. Only 6.8% sat in the volatile 30% to 70% middle, which leaves 77.5% of cells strictly always or never.
Whether an engine names you on a category question is closer to a fact about you than a coin flip. For most brands in that cohort the fact was no.
45.5% of mentioned cells flipped between positive and negative
Now hold the cells still and change the question. Instead of whether the answer named you, ask what it said. Kumar filtered to cells with at least three mentions and classified each run's sentiment.
- Always positive: 29.9%
- A mix of positive and neutral: 21.8%
- Always neutral: 2.3%
- A mix of negative and neutral: 0.5%
- Always negative: 0.0%
- Flipping between positive and negative: 45.5%
The stable signal is the one everybody reports
45.5% against 6.8% is where the 6.7 times comes from. Same cells, same runs. The number on the dashboard is the steady one, and the number nobody reports swings on nearly half the cells that have anything to report.
There's a quieter finding in that table we like more than the headline. Not one cell was consistently negative. Zero. The engines here never settled into a negative reading of a tracked brand, and negativity, where it showed up, passed. If someone is selling you AI reputation management on the theory that the models have turned against you, that row is worth keeping.
Mentioned and not recommended is now a named failure mode
Another group has come at the same gap from the other side. Will Jack, Noah Lehman, Keller Maloney and Sarah Xu, at Unusual, audited about 37,000 production runs over 215 commercially framed prompts across 19 sectors, against a 533-brand catalogue sorted into five prominence tiers. Unusual sells into this market too.
They split the problem into three stages. The model has to retrieve a page that references you, carry you out of the retrieval pool and into the answer text, then endorse you in the recommendation set instead of listing you as the option it passed over.
Category leaders turned up in nearly every relevant retrieval and won 25% to 41% of the slots they reached. Being in the room was solved for them. Getting picked wasn't. Further down the ladder it inverts, and 48% to 52% of the smallest brands never surfaced in any of the 37,000 runs.
In a companion paper the same team took 7,763 cases where neither ChatGPT nor Claude recommended a brand and sorted each into one of those three failures: never reached the model, reached it and wasn't mentioned, mentioned and wasn't recommended. The two providers gave the same diagnosis 95.1% of the time, with a clustered 95% interval of 94.3% to 95.7%. Agreement climbed as brands got smaller, from 81% on category leaders to 99.6% on long-tail regional brands. The picks diverge across providers. The reason you lost mostly doesn't.
A sentiment number needs a far bigger denominator than a mention number
The practical consequence is arithmetic. Kumar's own operating rule is that mention rates need at least three runs, and sentiment-weighted scores need at least ten prompts per platform per brand before they hold still. Sentiment reported off a handful of checks is reporting the sampler's luck.
This is why every visibility percentage we publish carries a Wilson interval, and why our alerting fires on the interval instead of the point estimate. Run a sentiment series at the same depth as a mention series and it will trip an alert something like seven times as often, and almost all of those will be noise.
- Keep mention and framing as separate series with separate sample sizes. One headline number averages a stable thing with a volatile one.
- When you lose, ask which of the three failures it was. The fix for never being retrieved has nothing in common with the fix for being the runner-up.
- Per engine. Cross-provider picks disagreed roughly two thirds of the time in the Unusual data.
- Set the alert threshold from the interval, so a wobble on nine samples doesn't page anyone.
What we do about it
Proofsource tracks sentiment and positioning next to prompt tracking and competitor share of voice, scores every visibility figure with a Wilson interval, and gates alerts on that interval. When you're named and passed over, the report's job is to say why they won, with the evidence attached.
Our own sampling is smaller and points somewhere slightly different. In the one category we tested we measured 9% to 27% day-over-day churn by engine in which names appear, more movement than a 6.8% flip rate suggests. Different cohort, different definition of a run, and we won't pretend those reconcile neatly.
The part we'd defend either way is the instruction that follows. Stop reading a mention count as a result. It's the entry condition.
Where this could be wrong
Both groups sell tools in this market and both are publishing their own production data. That isn't disqualifying, and we're in the same position when we publish ours. Neither one is an independent audit, so don't read them as one.
Kumar names the biggest limit himself. Sentiment is classified by a model over the spans that mention the brand, so the 45.5% mixes real variance in the answers with classifier noise, and the paper doesn't separate them. Treat 6.7 times as a ceiling on how unstable framing is, not a measurement of it.
That cohort is convenience-sampled and skews to SaaS, retail execution, fintech and Indian DTC. The Tier 1 group is 11 brands, and the paper flags its own small cells.
Unusual's failure-mode labels are model-assigned too, and their prominence ladder is hand-coded from public signals. It proxies how well known a brand is, not how big it is.
The windows are March to May 2026 and May 2026. Engines ship constantly, so take the shape and hold the decimals loosely.
Zero consistently negative cells is a statement about tracked brands on category prompts in one cohort. It isn't a claim that models never say anything bad about anyone.