How Proofsource measures AI visibility.
Proofsource measures AI visibility as the share of real AI answers that name your brand, asked the way your buyers ask, on ChatGPT, Perplexity, Claude and Google AI Overviews. Every percentage carries a 95% Wilson interval, every number opens to the answers behind it, and figures we cannot yet back with enough evidence are withheld, not estimated.
2 free scans, today and tomorrow: 25 questions on 4 engines. No card.
Three rules behind every number
Wilson interval on every percentage, so you can see how far the true rate could sit from the measured one.
Wilson (1927), the standard interval for a proportionengines asked the same questions: ChatGPT, Perplexity, Claude and Google AI Overviews.
Every plan, including the free trialestimated accuracy figures. A classifier score is published only when enough human labels back it.
Withheld below the minimum label countWhat does Proofsource count?
- Mention rate: the share of answers to a question that name your brand, per engine and overall.
- Who is named instead: every other brand in the same answers, ranked by how often each appears.
- Position: where in the answer your brand appears when it is named.
- Citations: the web pages each engine lists as sources, and which of them mention you.
Every answer is kept with its full text and sources, so each figure can be traced back to the answers it counts. The terms are defined in the AI visibility glossary.
How are the questions asked?
You approve a set of buyer questions: unbranded category questions such as "what is the best tool for X", plus a few questions that name your brand. Each one is asked on every engine in the panel, written the same way on each, so the engines can be compared.
On paid plans every question runs once per engine every day. The free trial runs 2 scans, today and tomorrow, with 25 questions on each engine. Unbranded questions are the real test: a question that names your brand will mention it whether or not the engine knows who you are.
How sure is each number?
An AI engine can give two different answers to the same question a minute apart, so one answer is a sample, not a fact. A percentage built from a handful of answers can sit a long way from the true rate.
Proofsource puts a 95% Wilson score interval on every percentage. Named in 6 of 20 answers is 30%, with an interval of about 14.5% to 51.9%. Named in 60 of 200 is also 30%, but the interval narrows to about 24.1% to 36.7%. Daily sampling is what narrows it, and a change is flagged only once it clears the interval.
The Wilson interval was introduced in 1927 and is the standard choice for a proportion from a small sample because it stays honest near 0% and 100% (reference).
How are the headline scores calculated?
Each report shows one score per area, from 0 to 100. Each is a weighted mean of measured parts. A part we could not measure is left out and named, never scored as zero. With too little evidence the score reads “not enough data”.
Technical score
Formula. round( sum(weight x area score) / sum(weight of the areas measured) ), each area score = round(100 x (passed + 0.5 x warnings) / determined checks)
- Crawl and indexing (weight 25). Area score from the site audit. The gate: a page that cannot be fetched, canonicalised or indexed cannot rank or be cited.
- Links and architecture (weight 10). Area score from the site audit. Internal links decide which pages crawlers reach; a broken link costs less than an unindexable page.
- Structured data and sharing (weight 15). Area score from the site audit. How machines read the page without guessing.
- Performance (weight 15). Area score from the site audit. Speed and Core Web Vitals affect ranking and abandonment directly, and are measured.
- Security and trust (weight 10). Area score from the site audit. A trust baseline most sites pass, so it separates few sites.
- On-page and content (weight 15). Area score from the site audit. Titles, descriptions, headings and body content are what a page is matched on.
- AI search access (weight 10). Area score from the site audit. Whether AI crawlers may read the site and llms.txt exists.
Missing data. An area with fewer than 3 determined checks is left out and the weights renormalised. Below 60 of 100 weight points measured, or without crawl and indexing, the report says not enough data and shows no score.
Uncertainty. An audit is a census of the crawled pages, not a sample, so there is no sampling interval. The range is a deterministic bound: the score if every undetermined check had failed, and if every one had passed.
SEO score
Formula. round( sum(weight x part score) / sum(weight of the parts measured) ), measured parts only
- Search visibility: keywords ranked (weight 10). log10(1 + keywords ranked), 10,000 keywords earns 100. How widely the site is found; log scale.
- Search visibility: rank buckets (weight 15). Mean click weight of the ranked keywords by position bucket, 0.25 earns 100. How well the site ranks where it is found.
- Search visibility: estimated monthly visits (weight 20). log10(1 + estimated visits), 100,000 earns 100. The outcome of ranking; a vendor estimate.
- Authority: referring domains vs named competitors (weight 25). 50 x brand referring domains / median of named competitors, capped at 100. Links are the strongest off-page signal, judged against the competitors the customer named.
- On-page (weight 20). The site audit's on-page and content area score. What the site controls about how a page is matched to a query.
- Content quality (weight 10). Mean verified score from the page-quality judge, only when a complete run exists. An opt-in LLM judgement on a sample of pages, the least deterministic input.
Missing data. A part that was not measured is left out, never guessed, and the rest renormalised. Without a measured search-visibility part, or below 50 of 100 weight points, the report says not enough data.
Uncertainty. Rankings, traffic and backlink counts are point-in-time vendor estimates, so no sampling interval is computed or implied. Each part shows its source and as-of date.
AEO score
Formula. round( (35 x mention + 25 x citation + 25 x share of voice + 15 x position) / 100 ), each part on 0-100, measured parts only
- Mention rate (weight 35). Share of answers naming the brand, 50% earns 100. The first thing a buyer needs; every other component requires it.
- Citation rate (weight 25). Share of answers citing the brand's site, 40% earns 100. The measurable step from visibility to traffic.
- Share of voice (top-10 frame) (weight 25). Brand share among the ten most-named vendors, 25% earns 100. How the brand compares with the vendors named beside it.
- Average position when named (weight 15). 100 x (20 - place) / 19, #1 earns 100. An earlier place is read first; weighted least because it rests on the fewest answers.
Missing data. Fewer than 10 answers: not enough data, no score. Share of voice needs 10 vendor mentions and position needs 5 named answers with a place, otherwise that part is left out and the rest renormalised.
Uncertainty. Answers are sampled, so the score is shown with a 95% range from the Wilson intervals of its rates (Monte-Carlo, fixed seed, conservative dependence), plus the score under five alternative weightings so the effect of the weights is visible.
How accurate is the brand-mention classifier?
Precision, recall and accuracy for the classifier that decides whether an answer names your brand, per engine, measured against human labels. Each label points at a real captured answer, and the classifier's current verdict is re-joined at report time.
Per-engine reliability figures are not published yet. We withhold them until the labelled set is large enough to report an honest number.
Precision and recall are measured against human-labelled gold sets, stratified per engine, topic, and brand. Each label references a real captured answer; the classifier's CURRENT verdict is re-joined at report time, so a model or prompt change is reflected without re-labelling. Figures below the minimum label count are withheld.
These figures are regenerated from the latest labelled run before each deploy.
How much do answers move on their own?
The noise floor is the natural run-to-run spread of a metric, measured on stored samples. A figure appears here only with the number of observations it rests on and the date it was measured.
Not yet measured. No noise-floor figure is published until one rests on a stored measurement.
Which model versions are sampled?
The exact model each engine was sampled with, and when we first saw it. A provider's silent model swap can shift figures before the next calibration run records it.
Model-version history is not published on this page yet.
What are the limits of this method?
- Silver labels bootstrap agreement but are not human-verified truth; only the gold-set figures on this page are.
- Sentiment is a 3-class agreement figure, not a per-class precision/recall.
- Engines are sampled with the model versions listed; a provider's silent model swap can shift figures before the next calibration run records it.
- Time-to-citation is measured only for pages Proofsource controls (unique-token probes), not for arbitrary customer pages.
What people ask about how Proofsource measures
How does Proofsource measure AI visibility?
It asks ChatGPT, Perplexity, Claude and Google AI Overviews the questions your buyers ask, reads every answer, and records whether your brand is named, next to whom, and which pages are cited. Mention rate is the share of answers that name you, and each rate carries a 95% Wilson interval.
How often are the questions asked?
Every day on paid plans, once per question per engine. The free trial runs 2 scans, today and tomorrow, with 25 questions on each engine.
Why put an interval on every number?
Because the same question can get a different answer a minute later. A mention rate from 20 answers could be far from the true rate; the interval shows how far. Proofsource only flags a change once it clears the interval, so a noisy day is not reported as a trend.
How accurate is the brand-mention classifier?
Its precision and recall are measured against human-labelled answers, per engine, and published on this page once the labelled set is large enough to be honest. Until then the figures are withheld rather than estimated.
Can I check the answers behind a number?
Yes. Every figure opens to the answers it was computed from, with the full text and the sources each engine cited, so you can recount it yourself.
See your own numbers, with their intervals.
See which buyer questions name your brand, who gets named instead, and what to fix. 2 free scans, today and tomorrow: 25 questions on 4 engines. No card.
Last updated