Compare

Proofsource vs LLMrefs

Our column is what we've built and measured. Where we haven't verified something about LLMrefs, this page says so instead of filling in the cell.

See every tool

2 free scans, today and tomorrow. No card · no sales call

Their side · unverified by us

Everything this page says about LLMrefs came from LLMrefs.

We haven't published a teardown of LLMrefs, and we're not going to describe a product we haven't run. Their site is the current answer on what they do and what it costs.

We haven't run LLMrefs against a domain of our own, so no number on this page is theirs. Everything below is ours, and you can check all of it with a free trial.

Side by side

8 rows, and we only filled in the ones we can be held to.

The Proofsource column is what you get today. The LLMrefs column gets a specific only where we've verified it ourselves.

QuestionProofsourceLLMrefs
Price to start2 free scans: 25 questions on 4 engines, today and tomorrow, no cardCheck their site[4]
Time to first resultThe same day you sign upCheck their site[4]
Sales processFree trial with no sales call; Custom through Contact sales in the appCheck their site[4]
Statistical confidence95% Wilson interval on every number; alerts wait until a change clears it[1]Check their site[4]
Engine coverageChatGPT, Perplexity, Claude and Google AI Overviews on every plan; Grok and DeepSeek as optional add-ons on CustomCheck their site[4]
Re-read cadenceDaily re-runs on Custom, because answer sets move 9-27% day over day in our sampling; the free trial runs 2 scans a day apart[2]Check their site[4]
Root cause and fixEvery gap carries an ID, the pages winning that phrasing, and a draft held for approvalCheck their site[4]
Getting your data outFull history as CSV or JSON, any date rangeCheck their site[4]
Verdict

Every “Check their site” in that column is a cell we could have guessed at and didn't.

Our side · built

4 differences that hold whoever's in the other column.

Each one is a mechanism, so you get what runs, how often, and what would catch us out if it weren't true.

01

The free trial is the real report.

The free trial asks 25 questions on ChatGPT, Perplexity, Claude and Google AI Overviews, today and tomorrow, and hands back what a subscriber reads: your index score on the same axis as the category leader, the coverage matrix of phrasings that do and don't return you, and the gap register. No card, no call, no gate. One per company domain.

02

Every number arrives with the interval it was measured at.

Ask an engine the same question twice and you'll get two answers, so one run is a draw and not a fact. We sample each prompt repeatedly, publish a 95% Wilson interval on every score, and hold alerts until a change clears it. A tool that treats a single run as a finding will sell you progress on a day you did nothing.[1]

03

Four engines answer the same prompts, on every plan.

ChatGPT, Perplexity, Claude and Google AI Overviews run side by side from the free trial upward. On Custom the sampler re-runs daily on its own, not when someone remembers to look. The full history leaves as CSV or JSON over any date range, so the record is yours.

04

Each gap names the page that beat you, then drafts the answer.

A gap gets a stable ID, a severity, the sources winning that phrasing, and what those pages have that yours doesn't. Then a draft built on exactly that, held at a human approval gate, because nothing should publish itself in your name. A score with no root cause leaves you to do the diagnosis.

How to judge any of them · measured

Any tool in this category can show you a score. Ask it these three questions instead.

Every figure below is our own live measurement, scoped to the engines and category we sampled. Try the questions on us first.

Test 1

Ask what it does when tomorrow's answer is different.

Answer sets move 9-27% day over day in our own sampling. Perplexity moved most at 27%, Claude least at 9%. Read the engines once and what you've got is a document: accurate the morning it was written, quietly out of date by the time you present it.[2]

Test 2

Ask it to name the pages the model is quoting.

In the category we tested, listicles were 44% of what ChatGPT cited for “best X” questions. Review sites 24%, vendor pages 18%, editorial 14%. Most of those listicles take submissions. If a tool can't name them, it can't tell you which gap is the cheapest to close.[3]

Test 3

Ask what it tells you to skip.

No engine we track documents llms.txt as a ranking signal, so we don't sell it as a fix. Any list of recommendations with nothing struck off it is selling you work.

Verdict

A shortlist of five or six names holds across almost every phrasing we test. It's stable enough to be worth getting into, and stable enough to stay in once you're there. So whatever tool you pick has to keep reading after the first report.

Questions

Proofsource or LLMrefs: what buyers ask

What is the difference between Proofsource and LLMrefs?

We haven't published a teardown of LLMrefs, so its own site is the current answer on what it does and what it costs. Proofsource is an AI visibility platform: it asks ChatGPT, Perplexity, Claude and Google AI Overviews the questions your buyers ask, puts a 95% Wilson interval on every number, and names the pages behind each gap with a draft fix held for your approval.

Is there a free trial of Proofsource?

Yes. 2 free scans: 25 questions on ChatGPT, Perplexity, Claude and Google AI Overviews, asked today and tomorrow, with no card and no sales call. One free trial per company domain, and the report is about your own domain.

Which AI engines does Proofsource track?

ChatGPT, Perplexity, Claude and Google AI Overviews on every plan, with Grok and DeepSeek as optional add-ons on Custom.

How often does Proofsource re-read the answers?

Every day on Custom, because answer sets move 9-27% day over day in our own sampling. The free trial runs 2 scans a day apart.

Why does the LLMrefs column say “Check their site”?

Because we haven't verified that cell ourselves, and a guessed cell only tells you what the author wanted to be true. Where we decline to answer, LLMrefs's own pricing and docs are the place to look. If you find a row we got wrong, tell us and we'll change it.

Sources

Where each claim on this page comes from.

  1. Wilson score interval. The interval we put on every visibility percentage. Introduced by E. B. Wilson in 1927 and standard for binomial proportions, because it stays honest at small sample sizes and near 0% or 100% where the naive normal approximation breaks.
  2. Proofsource sampling: day-over-day change in which brands get named, same question one day apart. Perplexity 27%, ChatGPT 18%, Google AI Overviews 12%, Claude 9%. How we measure.
  3. Proofsource measurement of the sources ChatGPT cited for “best X” questions, in the one category we tested: listicles 44%, review sites 24%, vendor pages 18%, editorial 14%.
  4. LLMrefs's own site, pricing and docs, for every cell marked “Check their site”. We haven't run LLMrefs ourselves. Check a row and find we've got it wrong? Tell us and we'll change it here.

Every comparison we've written

A free trial settles this faster than either of us can write about it.

25 questions on ChatGPT, Perplexity, Claude and Google AI Overviews, asked today and tomorrow, no card, and the report is about your domain instead of ours: your index score on the same axis as the brand ahead of you, and the phrasings where a buyer never reaches you.

Last updated