Handed the exact paragraph, eight tools named the wrong source 60% of the time
Here's a test that should be easy. Take a paragraph out of a published article, paste it into an AI search tool, and ask which article it came from, who published it, and what the URL is. The text is sitting right there in the prompt. Nothing has to be guessed.
The Tow Center for Digital Journalism ran exactly that in February 2025 across eight tools: ChatGPT Search, Perplexity, Perplexity Pro, DeepSeek Search, Copilot, Grok 2, Grok 3 in beta, and Gemini. Twenty publishers, ten articles each, 1,600 queries in total. The excerpts were picked so that pasting them into ordinary Google search returned the original inside the first three results.
Collectively the tools answered more than 60% of the queries incorrectly. Perplexity was the best of the group and still got 37% wrong. Grok 3 got 94% wrong. DeepSeek attributed the excerpt to the wrong source 115 times out of 200.
The confidence is the part worth sitting with. ChatGPT misidentified 134 articles and signalled any uncertainty fifteen times across two hundred responses. It never once declined to answer.
Perplexity's own docs say the markers are prompt-dependent and the result list is the truth
That reads like a model quality problem. It's closer to an architecture problem, and you can read the architecture in the vendors' own developer documentation.
A web search through Perplexity's Agent API comes back as two separate objects. One is a search_results item carrying the queries that ran and every result the tool pulled, each with an id, a url, a title and a snippet. The other is the assistant message, which is the answer text. The bracketed markers everybody treats as citations live inside that answer text, and Perplexity's reference page for the tool says whether the model adds them at all is prompt-dependent, so you should ask for them explicitly. Then it tells developers what to rely on instead: regardless of markers, treat the id and url fields of each search_results entry as the source of truth for citations.
Read that with a brand's eyes. The vendor is telling its own developers that the visible citation is not authoritative and the retrieved set is. Anyone measuring AI visibility from outside only ever sees the visible citation.
The same documentation warns against asking a model to emit URLs inside structured output at all, because a model writing links can produce malformed or fabricated ones, and points developers back at the search results. The Tow Center measured the consumer-side version of that failure. More than half the responses from Gemini and Grok 3 carried fabricated or broken URLs, and 154 of Grok 3's 200 citations landed on error pages.
A citation the system extracted carries character offsets
Put that next to a citation the system computes instead of the model writing it. Anthropic's Citations feature on the Claude API returns, for each claim in the answer, a block holding the exact cited_text, the index of the document it came from, and a start and end character offset into that document. The docs are explicit about why the shape matters: because the API parses citations and extracts cited_text directly, citations are guaranteed to contain valid pointers to the provided documents.
Google's grounding API for Gemini is built along the same lines. A grounded answer comes back with the searches the model actually executed, plus url_citation annotations that each carry a start_index and an end_index into the answer text, so you can highlight the exact span and name the URL behind it.
The machinery exists, then. A citation that's a pointer the system computed is checkable. A citation that's a string the model wrote is a claim, and it deserves the same scepticism as every other claim in the answer. The surfaces buyers actually use mostly show the second kind.
Your paragraph on somebody else's domain collects the credit
The Tow Center found one failure that lands harder on brands than on anyone else. The tools repeatedly cited syndicated or republished copies instead of the original. Perplexity Pro cited republished versions of Texas Tribune articles in three of ten queries, and Perplexity has a partnership with the Texas Tribune. ChatGPT cited a Yahoo News republication of a USA Today article.
A brand's version of this is duller and far more common. Your comparison page explains the real difference between two plan tiers. A roundup blog paraphrases it. When somebody asks the category question, the roundup takes the link, your wording is in the answer, and your domain isn't on the page anywhere.
In our own sampling, on a best-in-category question in the one category we tested, listicles and roundups accounted for around 44% of the sources ChatGPT cited. That's ours and it's scoped to what we sampled. It fits the shape of the problem though: citation slots are a small, contested surface, and third-party pages are very good at occupying them.
Count three findings instead of one
If the link list is an unreliable readout of what the engine used, then a single citation count is the wrong instrument. Three questions come apart cleanly, and each one leads somewhere different.
- Did the answer name you at all? Presence in the prose is its own measurement, and it moves independently of whether a link showed up.
- Did the answer credit you, or credit a page carrying your material? If it's the second, you have a placement problem, and rewriting your own page won't touch it.
- Could the engine reach you at all? A missing citation has a boring explanation a good deal of the time, and a bot accessibility check settles that before anyone rewrites anything.
What we do about it
This is why Proofsource runs citation-gap and ghost-citation detection instead of reporting one number, and why every data point is labelled by provenance: observed in an answer, or generated. An answer that quotes you without linking you and an answer that never saw you look identical on a dashboard that counts links.
It's also why the sampling records the answer text, not only the source list. If your phrasing shows up in the prose while somebody else's domain shows up in the sources, that's a finding, and it's a different job from not being retrieved.
Where this could be wrong
The Tow Center tests ran in February 2025 and published that March. That's eighteen months ago and every product in the set has shipped many versions since. Take the 60% as evidence that this class of failure is real and measurable, not as today's error rate. We have not re-run it.
The study's own limitations section says each query was run once, and that anyone re-running the same prompts would very likely get different outputs. Attribution accuracy is itself a noisy measurement.
Pasting an excerpt and asking who wrote it is a provenance probe, not how a buyer shops. It stresses attribution directly, which is the point of it, but it doesn't tell you how often attribution goes wrong inside a normal category question.
The API behaviour above is the developer surface. ChatGPT, AI Overviews and Perplexity's own website are not guaranteed to run the same code path, and we're reasoning from documented developer behaviour rather than from inside the consumer products.
And the honest limit on our own advice: from the outside you cannot always separate a page that was retrieved and not credited from a page that was never retrieved. Checking crawler access narrows it. Looking for your own phrasing in the answer text narrows it further. Neither closes it.