One in four sources AI shows you for a “best product” question has a commercial interest.
A study of 106,758 citations across ChatGPT, Google AI Overview and Perplexity.
Ask ChatGPT, Google or Perplexity for the best VPN, the best mattress or the best pet insurance and you get a confident shortlist. In collaboration with DataPulse Research, we traced the sources behind 3,334 such bottom-of-the-funnel questions in English to see what those shortlists are built from. One in four has a commercial interest. Most are affiliate reviews, a publisher’s own test that earns a commission on a sale; paid rankings, written or bought by a vendor, are a much smaller share. What you ask matters more than which tool you ask: gadget questions draw three times as many commercial sources as education questions. And ten publishers supply a quarter of them.
- 25.4% of the sources ChatGPT, Google AI Overview and Perplexity show for English “best X” questions have a commercial interest, properly disclosed on the page. Most are affiliate reviews. Explicit “sponsored” labels are rare in English, less so in German. The 25% is a conservative reading: disclosures that no scanner reads, and pages with none, are not in it.
- Nearly nine answers in ten (86%) use at least one such source, ChatGPT most often (91%) because it cites the most pages. Per citation the order reverses: Perplexity 31.1%, ChatGPT 22.5%.
- The category matters more than the tool. Consumer electronics questions draw 41% of their citations from commercial pages, education questions 13%.
- Ten publishers, review brands such as TechRadar and PCMag and commerce sections such as Forbes Advisor, account for a quarter of all commercially disclosed citations.
- Paid rankings on major news sites are cited too, from the Los Angeles Times to the Houston Press.

- How the study was done
- A quarter of the sources AI shows you have a commercial interest
- Paid rankings on major news sites are cited too
- Nearly nine answers in ten use at least one commercial source
- The category matters more than the tool
- Ten publishers supply a quarter of the commercial layer
- Why 25% is a conservative reading
- What this means
- Method and limits
How the study was done
We wrote 3,334 English comparison questions of the kind people ask an AI assistant before buying something: “best 4K TV compared: which one wins?”, “which payroll software should I choose?”, “what pet insurance do experts recommend?”. They cover 748 product and service types, grouped into 15 sectors. The data comes from BuzzView, an AI-search analytics platform, which ran every question through ChatGPT, Google AI Overview and Perplexity in April 2026 and recorded every URL the three tools cited.
We then loaded each cited page in a headless Chromium browser and checked it for the visible markers of commercial content: affiliate disclosures (“we may earn a commission”), partner-content notices (“in partnership with”), and explicit advertising labels (“sponsored”, “paid post”). The same questions were also run in German; this study reports the English results. Full method and limits are at the end.
A quarter of the sources AI shows you have a commercial interest
Share of all English citations by marker type. Markers overlap, so the bars do not sum to 25.4%.
The commercial layer is mostly affiliate content: review and comparison pages that earn a commission when a reader buys through their links, and say so. Explicit advertising labels, the “sponsored” or “paid post” kind, are on under 2% of citations. That layer is filtered by context rather than validated page by page, and spot checks show it still catches phrases such as “employer-sponsored plans”, so read it as an upper bound on an already small number. The same questions run in German return 29%, on a German marker set and a largely different set of publishers.
Paid rankings on major news sites are cited too
Affiliate reviews are the bulk of that quarter. Paid content is the smaller, sharper case: rankings written or paid for by a vendor and published on a newspaper’s site with a label. We re-ran questions from the corpus on 7 September 2026, three runs per question and tool, and captured the pairs below: left, the publisher page with its label; right, an AI answer citing it. The red box marks the label on the page and the source in the answer.
“Best online psychic reading sites”
“Best hair transplant clinic in Turkey”
“What are the best debt relief companies?”
“What is the best pet insurance in the US?”
“What is the best collagen supplement in the UK?”
Nearly nine answers in ten use at least one commercial source
Back to the corpus. Citation shares understate what a single answer contains, because one answer draws on many sources. Counting only sources we could match to the scan, ChatGPT averages 19 per answer, Perplexity 10, Google AI Overview 9.
| Citations | Commercial share of citations | Answers with ≥1 commercial source | Answers mostly built on them | |
|---|---|---|---|---|
| 33,155 | 31.1% | 89% | 17% | |
| 21,028 | 23.6% | 77% | 10% | |
| 52,575 | 22.5% | 91% | 5% | |
| All three | 106,758 | 25.4% | 86% | 11% |
English answers with at least one source that could be matched to the scan (n = 8,536). “Mostly” means more than half of an answer’s matched sources carried a commercial disclosure. The pooled row is citation-weighted, so ChatGPT, which cites the most pages, carries about half of it.
ChatGPT has the lowest share per citation but cites twice as many pages, so it has the highest chance of including at least one. On Perplexity, one answer in six builds its recommendation mostly on commercial sources. The spread between tools, eight points, is a fraction of the spread between categories, 29 points.
The category matters more than the tool
The share of commercial sources varies far more by what you ask than by which tool you ask. Questions about gadgets, sports gear and home appliances are answered largely from affiliate review sites. Questions about clinics, lawyers and courses draw on provider pages and directories instead.
Share of citations that lead to commercially disclosed pages, by sector. 3,320 English questions mapped to 15 named sectors by product type, plus a residual group of 796 citations not shown; 14 questions could not be assigned.
At product level the spread is wider still. Among the 459 product types with at least 80 citations, the highest are Cordless vacuum (61%, 120 citations) and Fitness tracker (60%, 121). At the other end, 3 of 165 citations for divorce mediation services and 3 of 159 for hair restoration clinics carried a marker. The low end is not neutral ground: those are the sectors where the sources are mostly provider pages and directories, which carry nothing a scanner can read as a disclosure.
Ten publishers supply a quarter of the commercial layer
The commercial sources are concentrated. Ten publishers account for 6,786 of roughly 27,100 commercially disclosed citations, a quarter of the total. Most are affiliate-supported review brands whose model is disclosed on every page: TechRadar’s “when you purchase through links on our site, we may earn an affiliate commission” sits under every article, which is why nearly all of its citations count as commercial. The rest are the commerce divisions of general-news brands, such as Forbes Advisor and the New York Times’ Wirecutter, which the study counts by URL path as hand-verified commerce sections rather than by a detected label.
Number of citations, all three tools, English, April 2026, that lead to a page carrying an affiliate, partner or advertising disclosure, or sitting in a hand-verified commerce section (Forbes Advisor, NYT Wirecutter).
The disclosure says a commission may be earned. It does not say which brands pay, how much, or whether it moves the ranking, and neither does this study. What the data does show is how much of the recommendation layer these ten carry: across the corpus each is cited hundreds to thousands of times, and in consumer electronics the affiliate review brands dominate the source list.
Why 25% is a conservative reading
Every figure above rests on pages that say what they are, in a form a scanner can read. Two kinds of page escape that, and the gap is in our measurement, not in the pages.
The first is a disclosure that is not a label. The Jerusalem Post page below ranks hair-transplant clinics. There is no “sponsored” or “paid” badge on it. The commercial relationship is stated in the byline: “By in cooperation with Dr. Terziler Exclusive Clinic”. The page discloses the relationship to its readers. The only finding here is about our scanner: a label detector cannot see a disclosure written as a byline, so pages like this count as non-commercial in every figure in this study.
“Best hair transplant clinic in Turkey”
The second is no disclosure at all. None of the pages shown in this study is of that kind. The study cannot see a page that carries none: it counts as non-commercial here whatever its actual status. That matters most in the sectors with the lowest measured share, clinics at around 2%, divorce services at 2 to 3%, heat-pump installers at 4%, where the sources are mostly provider pages and directories. There the measurement is least able to tell neutral sourcing from commercial content that simply is not labelled.
One more observation from the live checks. We ran twenty further questions in categories known for paid rankings, from CBD gummies to gold IRAs, on all three tools from a US IP, and checked every cited page for a sponsored, paid or partner label. Exactly one new labelled page surfaced, the LA Times debt-relief page above. The thinness is itself the finding. Explicit paid-post labels on large English news sites are rare. The commercial layer AI tools cite is overwhelmingly affiliate content, and one of the rankings we found carries its disclosure in a byline rather than a label, where no scanner would catch it. The measured 25% is therefore a conservative reading, and the unmeasured part is largest in exactly the sectors where the measured share is lowest.
What this means
When an AI tool names the best product in a category, it is summarising a handful of review and comparison pages, and a quarter of the sources behind that summary have a commercial interest. The affiliate line, the paid-program bar and the cooperation byline all stay on the source page; what reaches the reader is a source chip with the publisher’s name. The recommendation is judged on the source’s authority, not on its commercial model, and the context is one click away.
Method and limits
- Prompts. 6,668 comparison questions, 3,334 in English and 3,334 in German, generated by applying a template set (“best X compared”, “which X should I choose”, “top 10 X”, “what X do experts recommend” and others) to a list of 748 product and service types. Sectors were assigned by keyword rules on the product type; the mapping is available on request. 14 English prompts could not be assigned and are excluded from the sector chart.
- Citations. Collected through BuzzView in April 2026 on ChatGPT, Google AI Overview and Perplexity, with every English project set to English and United States. 139,033 unique URLs in total. English answers: 3,333 from ChatGPT, 3,334 from Perplexity, and 2,448 from Google, where an AI Overview rendered for 73% of questions and a record exists only for those. 106,758 English citations from 8,536 answers could be matched to a scanned page; one ChatGPT answer in six (16.9%) returned no matchable source and is excluded from the per-answer figures. The corpus is an April snapshot; the examples were re-captured in September and reproduced on the tools stated.
- Scan. Every URL was loaded in headless Chromium (Playwright) with full JavaScript rendering; 132,335 (95.2%) loaded successfully. The scanner checked rendered text and source for 22 markers in three layers (advertising labels, affiliate disclosures, partner-content notices) plus a path-based check for hand-verified commerce sections (Forbes Advisor, Wirecutter, CNN Underscored, NBC Select, IndyBest, USA Today Blueprint, Rolling Stone product recommendations).
- False positives. About 7,800 hits from cookie banners, navigation, product descriptions and running text (“employer-sponsored plans”) were removed by context rules; ambiguous labels were verified in the page structure. Precision was checked against a hand-verified reference set; affiliate disclosures scored over 95%. English advertising labels were filtered by context only and the 1.8% figure is an upper bound.
- Two denominators. Headline shares count every displayed citation, because a citation is what a reader is exposed to and a source cited twice shapes two answers. Per unique URL the English share is 20.5%. Residual false positives could lower the measured share by a few points; pages with no machine-readable disclosure push the true share the other way, by an amount the study cannot measure.
- What is not measured. Pages without a machine-readable disclosure count as non-commercial whatever their actual status. The shares apply to “best X” questions; other question types produce a different source mix. Blocked URLs (4.8%) may skew the result if they are more often commercial.
- Examples. Captured on 7 September 2026 with a scripted desktop browser from a US IP address (via VPN), Google set to hl=en and gl=us, Google and ChatGPT signed out, Perplexity signed in, three runs per question and tool. The debt-relief pair was captured on 14 September 2026 through a US proxy, signed out. Pages were chosen because their labels are unambiguous, not because they are typical, and AI answers change between runs and locations.
- Live sweep. On 7 September 2026, twenty further “best X” questions in categories known for paid rankings were run once each on all three tools from the same US IP. Every cited URL was checked automatically for sponsored, paid, partner and advertorial label strings and path segments, and against a list of known newspaper sponsored-content sections. One run per question, so absences are weak evidence.























