The organic answer is the part you cannot buy
OpenAI began testing ads in ChatGPT on 9 February 2026, and by its 11 August update the pilot had reached the United Kingdom, Mexico, Brazil, Japan and South Korea. Ads are labelled sponsored, sit separated from the organic answer, and in OpenAI's own words do not influence the answers it gives you.
So the recommendation is not for sale. Which leaves the question every challenger is already asking: if the model has heard of your competitor and not of you, is there a way in?
With identical specs, the recognised brand was picked in 100% of 670 trials
Xi Chu at Trine University and YuPeng Hou at Texas A&M measured it in a paper submitted in June 2026 and revised in August. They built sets of ten skincare products: one real brand, CeraVe or Paula's Choice or EltaMD, and nine fictional names screened to confirm the models had never heard of them.
In the first condition every product carried the same rating, price, review count and ingredient description. Only the name differed, so a fair model picks each about 10% of the time. Across 670 valid trials the real brand won every single one, in all three models and all four subcategories. Not one fictional brand was ever chosen.
That is the bleakest possible reading for a challenger, and the least useful, because the condition it describes is one you are probably reproducing by accident.
A 0.075-star edge flips it, and brand then explains 1.2% of the ranking
Give the unknown product better specs and the real one worse, and across 2,769 trials the models stayed loyal between 1.7% and 4.6% of the time. Shrink the advantage and the transition turns out to be a step rather than a slope. At identical specs the unknown brand won 3.6% to 6.0%. At the smallest advantage tested it won 64% to 80%, and piling on more bought little. Interpolated to a coin flip, across 9,220 trials, the threshold is small enough to write down.
- A 0.075-star rating advantage, less than the gap between 4.3 and 4.4 stars
- A 1.6 times higher review count
- A 7.3% lower price
Recognition is what the model falls back on when you say nothing
A variance decomposition over 14,395 trials settles what drove the ranking. Product parameters explained 82.4%, list position 6.5%, brand identity 1.2%. Brand had no measurable effect when quality was clearly high or clearly low. Its influence peaked in the middle, where the information was ambiguous. Ask this per engine rather than in aggregate: at the smallest rating advantage Claude flipped 11% of the time against 94% for GPT-4o-mini and 88% for Gemini.
The strongest lever they found was a clinical claim that did not exist
Holding every specification identical and changing only the copy, authority language broke the incumbent's hold 73.3% of the time and social proof 50.7%, against a baseline near 4%. Anchoring, scarcity and loss aversion barely registered, between 9.6% and 12.9%. The models ignored urgency and treated a citation-shaped sentence as evidence worth about 0.17 rating points, or a 15.3% discount.
Here is the part a summary will drop. Those stimuli were invented clinical trials and fake dermatologist endorsements. The authors call that potential false advertising and restrict their own recommendations to real certifications and published trials. The measured lever is fabrication, which is a finding about the models' evidence checking rather than a tactic.
When every brand optimises, the answer returns to the incumbent
The third experiment should change how you budget. Scaling the challengers writing authority copy from none to all nine, over 4,800 calls: with nobody optimising the incumbent survived 100% of the time, with one challenger 19.8%, with all nine 93.8%. The signal stopped being a signal once everybody carried it, and the model went back to the name it knew.
Per-brand payoff decayed from 0.802 for the first mover to 0.007 under universal adoption, a half life of roughly 1.4 competitors. Across 4,745 trials, challengers that did not optimise got zero recommendations. The authors call this a prisoner's dilemma and they are right. A tactic a competitor can copy in an afternoon has a payoff with a half life of about one competitor. A real difference you can prove does not.
In the real thing, the model has to find the fact first
Every condition above hands the model a clean specification table inside the prompt. Live engines assemble one, and that gap is where the practical work sits. In the category we sampled, listicles were 44% of what ChatGPT cited for a best-of question, review sites 24% and vendor pages 18%. Those are our figures, scoped to our samples, not a general law.
The shape is the point. If the roundups answering your category question list you with no rating, no review count and not the one specification you win on, then as far as the model is concerned you are in the identical-specs condition. You are not losing the comparison. You are not in one.
- Find the single dimension you genuinely beat the category leader on, and check it appears as a number in the third-party pages the engines cite, not only on your own site.
- Pull the answers where a competitor got the recommendation and read what the model said the difference was.
What we do about it
Proofsource samples the engines repeatedly rather than once, scores every visibility figure with a Wilson interval, gates alerts on that interval, maps which pages get cited beside you, and checks whether the crawlers can read your pages at all.
The boundary, stated plainly: we measure whether the engines name you, what they cite when they do, and whether AI crawlers and AI referral traffic reach your site. We do not measure whether the sale happened.
Where this could be wrong
One category, skincare, picked because buyers cannot judge quality before purchase. The search-goods replication covers the first experiment only, and the authors say they have not tested the language or multi-brand results outside it. Three closed models, one buyer persona, temperature 0.7, with a cluster bootstrap over 24 blocks that the main effects survive.
Nine of the ten products in every set were fictional. That isolates brand recognition cleanly and removes everything a real competitor has: reviews in the wild, third-party coverage, pages the model can retrieve. The 100% is not a real market.
Most of this tests what the models know rather than what they retrieve, and the engines buyers use are retrieval systems. The paper's own retrieval probe is minimal, one embedding model and no re-ranking, and the authors decline to generalise from it. It points the other way: the retrieval layer gave the real brand no credit for being known and its recommendation rate fell to 0%.
Our 44% is a single measurement on our own sample, published as ours and not as a property of the web.