Method5 min read

Ask for the best software and 73% of the cited sources are third-party sites

An answer is a set, and most of the set is written by people who aren't you. The pages cited beside you do more work than your own page does.

Seventy prompts pulled 1,702 citations, about eight per answer

Most conversations about AI visibility treat an answer as a verdict on one page. Yours. It reads that way from the inside, because you only ever check whether you're in it.

An audit published in September 2025 by Arlen Kumar and Leanid Palkhouski, at UC Berkeley and Wrodium Research, put numbers on the shape of the thing. They ran 70 product-intent prompts across 16 B2B SaaS verticals through Brave Summary, Google AI Overviews and Perplexity's sonar-pro, and collected 1,702 citations pointing at 1,100 unique URLs.

Seventy prompts through three engines is 210 answers. Divide 1,702 citations across them and you get roughly eight cited URLs per answer. That's the real unit. Eight slots, and seven of them belong to somebody else even on a good day.

On software questions, 72.7% of cited domains weren't the vendors'

A second study from the same month gets at who those neighbours are. Mahe Chen, Xiaoxuan Wang, Kaiwen Chen and Nick Koudas at the University of Toronto ran 1,000 consumer ranking prompts, ten categories at a hundred each, in the US and Canada against both Google and a web-enabled GPT. They took the top ten URLs per query and sorted every domain into Brand, Social or Earned, meaning independent media, review and comparison sites.

In Software Products in the US, AI search returned 72.7% Earned, 26.7% Brand and almost no Social. Google on the same intents came back 43.7% Brand, 45.4% Earned, 10.9% Social. Canada was wider: AI search gave 74.2% Earned and no Social at all, while Google leaned 53.8% Brand. Consumer electronics in the US went further still, to 92.1% Earned.

Read that as a brand and it stings a little. On a category question, most of what the model is working from was written by people you don't employ, on pages you can't edit.

Co-citation was defined in 1973 and the definition is the whole point

There's an older idea that fits this exactly. In 1973 Henry Small defined co-citation in the Journal of the American Society for Information Science as the frequency with which two documents are cited together, measured by comparing the lists of citing documents and counting identical entries. Clusters of co-cited papers, he showed, map the structure of a scientific specialty.

The property worth carrying over is the one Small built the measure on. Co-citation is created by whoever does the citing. Two papers become neighbours because a third party kept reaching for both of them. Neither can declare the relationship, and neither can opt out.

Transpose that to an answer engine and the consequence is uncomfortable. You don't choose the pages you appear beside. The engine assembles the set, fresh on each request, and the buyer reads all eight sources as one thing. Two vendors with near-identical pages can land in different neighbourhoods and get read differently for it.

Only 134 of 1,100 cited pages turned up on more than one engine

Neighbourhoods also don't transfer. In the Kumar and Palkhouski corpus, 134 URLs were counted as cross-engine citations. Against 1,100 audited URLs that's about one in eight, so roughly seven eighths of the pages cited were cited by a single engine. The 134 that did travel scored 71% higher on the paper's page-quality index than the single-engine ones, which is a decent argument for the boring structural work.

The practical read is that a place in Perplexity's source set tells you very little about Google AI Overviews. If your dashboard collapses engines into one visibility number, it's averaging over three mostly separate populations of pages.

Citing other people lifted the fifth-ranked site by 115%

Here's the part that changes what you do on Monday. The paper that named generative engine optimisation, by Pranjal Aggarwal and colleagues at Princeton and IIT Delhi, tested nine content rewrites against a 10,000-query benchmark spanning 25 domains.

Three methods came out on top: adding citations to credible sources, adding quotations, and adding statistics. Together they produced a relative improvement of 30-40% on the paper's position-adjusted word count metric and 15-30% on its subjective impression score.

Split by search rank, the Cite Sources method raised visibility by 115.1% for sites already ranked fifth, while the top-ranked site's visibility fell by an average of 30.3%. The lever there isn't authority. It's citability. A page that quotes credible sources hands the model something liftable, and it puts your sentence in the same paragraph as the sources you quoted.

Which is the co-citation point running the other way. Who you cite shapes who you get cited with.

Measure the neighbourhood, not only your own row

If the set is the unit, four questions are worth more than a citation count.

  • Which domains keep appearing together for your category. If the same four pages recur across your prompts, that group is your category's actual source list, and it's knowable.
  • Whether a competitor sits in a set you're missing from. That's a placement gap, and no amount of rewriting your own homepage closes it.
  • Per engine, kept separate. On the numbers above the sets barely overlap, so one engine's neighbourhood predicts another's poorly.
  • Which co-cited pages accept submissions or corrections. Most roundups do, and asking is usually the cheapest way into a set you're outside.

What we do about it

This is why Proofsource runs co-citation mapping next to prompt tracking and citation-gap detection. The report shows which sources travel together for your category and which ones you're missing from, and the client report carries a listicle source-mix and freshness view for the same reason.

Our own sampling points the same way. On a best-in-category question in the one category we tested, listicles and roundups were about 44% of the sources ChatGPT cited, review sites 24%, and vendor pages 18%. That's ours, one category and one engine, and we wouldn't defend it as a general law.

It also means we quite often tell people the cheapest fix sits on a page they don't own and we can't touch. We'd rather say that than sell a rewrite that won't move anything.

Where this could be wrong

Both audits ran in September 2025. That's a year ago, and every engine in them has shipped many versions since. Treat the direction as well evidenced and the exact percentages as dated.

The Kumar and Palkhouski study is observational, English-language, B2B SaaS, and collected at one point in time from 70 prompts. It queried Perplexity through the sonar-pro API, which need not behave like the consumer product. Their 134 figure is a count in that corpus, and they publish no pairwise overlap, so we can't say which two engines share most.

The Toronto study classified every domain using GPT-4o-search-preview. Model-assigned labels carry model errors, and the Brand versus Earned line genuinely blurs for a vendor-funded comparison site.

There's a disagreement we can't resolve here. Several widely-circulated trackers report community platforms such as Reddit taking a very large share of AI citations, while this study found Social close to zero. We haven't read those trackers back to a primary, so we're not quoting a number from them. We're flagging the conflict instead of picking a side.

The Aggarwal benchmark dates from 2023 and 2024 and was evaluated against engines two generations old. Take the mechanism from it and treat the 115% as an artefact of that setup.

Small's 1973 measure described a fixed printed record. An answer's source set is regenerated per request, so who you're cited with is a distribution. One answer establishes nothing about your neighbourhood.

Sources

Where the outside numbers come from

  1. AI Answer Engine Citation Behavior: An Empirical Analysis of the GEO-16 Framework · Arlen Kumar, Leanid Palkhouski, arXiv:2509.10762
  2. Generative Engine Optimization: How to Dominate AI Search · Mahe Chen, Xiaoxuan Wang, Kaiwen Chen, Nick Koudas, University of Toronto, arXiv:2509.08919
  3. GEO: Generative Engine Optimization · Pranjal Aggarwal et al., KDD 2024, arXiv:2311.09735
  4. Co-citation in the scientific literature: A new measure of the relationship between two documents · Henry Small, Journal of the American Society for Information Science, 1973
All notes

Questions

If most citations go to third parties, is our own site worth fixing?

Yes, for two reasons. Brand domains still took 26.7% of cited domains in the software category above, so your site is in the set, just outnumbered. And the structural work that gets a page cited at all, visible and machine-readable dates, clean semantic HTML and valid structured data, is the same work that got the cross-engine pages their higher scores. It's the floor. It just isn't the whole building.

How do we get into a co-cited set we're currently outside?

Find out which pages are in it first, then work out which of them are reachable. A surprising number of roundups, comparison pages and category directories take submissions or will fix an out-of-date entry if you write to them. That's a shorter path than rewriting your site and hoping. Where the set is all editorial coverage you can't request, the honest answer is that it's a slow problem.

Should we track each engine separately or roll them up?

Separately, then roll up if you want a headline. About one in eight cited URLs appeared on more than one engine in the corpus above, so a blended score is averaging across populations that barely intersect. A drop in the blend can be one engine moving while the others sit still, and you can't tell which from the blend.

Keep reading

Related notes

Your own numbers beat our best post.

The shortlist in your category is being written right now, whether anyone reads this page or not. A free trial, 25 questions on four engines today and tomorrow, tells you if your name is in it.