Six in ten sources the AI reads never reach an answer
Ten engines opened 11,162 different domains over three months. Only 4,476 of them ever showed up in an answer. The other 6,686 did the work and got none of the credit.
An AI answer has two layers. The links you can click, and everything the model read before it started writing. The second layer is much larger and nobody shows it to you.
We logged both layers for 20,879 answers. This article is about the gap between them.
What we measured
Every answer Bee LLM collected between 2 June and 22 August 2026, across all accounts:
| Answers analysed | 20,879 |
|---|---|
| Engines | 10 (ChatGPT web and API, Gemini, Perplexity, Claude, Copilot, AI Overviews, AI Mode, DeepSeek, Qwen) |
| Scope | 14 projects across 7 sectors, from accommodation to software |
| Links logged | 56,966 cited and 214,038 opened |
Two words that get used interchangeably and should not be:
- Cited. The AI leaves the page as a link inside its answer. The reader sees it and can click it.
- Opened. The AI read the page to build the answer and never showed it. No clicks, but it still shapes what the model ends up saying.
Six in ten domains never make it into an answer
Over the three months the engines opened 11,162 different domains. Only 4,476 were cited at least once. The remaining 6,686, or 59.9%, were read and never shown.
The per-answer figures point the same way. Each answer opens 10.3 pages on average and links 2.7.
A site can shape hundreds of answers and receive nothing back. No link, no visit, no line in anybody's report.
What the AI reads is not what the AI shows
Put the two layers side by side and the distortion is easy to see. On the left, the share of everything the engines opened. On the right, the share of everything they cited.
| Type of source | Share of what it reads | Share of what it cites |
|---|---|---|
| Company sites | 56.0% | 66.0% |
| Forums | 11.2% | 2.7% |
| Media and publishers | 9.5% | 11.2% |
| Reference (Wikipedia and similar) | 7.3% | 2.5% |
| Marketplaces | 5.0% | 5.9% |
| Institutions | 3.6% | 2.1% |
| Directories | 3.2% | 4.4% |
| Review sites | 1.4% | 1.9% |
| Social media | 1.3% | 2.4% |
The citation column leaves out the analysed brands' own sites and those of their direct competitors, so the figure is not inflated by self-citation. Counting those in, company sites would be 69.9% of all citations rather than 66.0%.
Two categories shrink on the way to the answer. Forums drop from 11.2% to 2.7%. Reference sites drop from 7.3% to 2.5%. Both are read far more than they are shown.
Company sites move the other way, from 56.0% to 66.0%. When the model decides what to put in front of a reader, it reaches for a page written by a business about its own product.
How much of each type survives the edit
The same thing again, this time as a conversion rate. How many of the pages opened in each category end up as a visible link.
| Type of source | Pages opened | Pages cited | Reaches the answer |
|---|---|---|---|
| Social media | 2,884 | 1,108 | 38.4% |
| Marketplaces | 10,675 | 3,949 | 37.0% |
| Company sites | 119,808 | 39,822 | 33.2% |
| Directories | 6,764 | 2,035 | 30.1% |
| Review sites | 3,094 | 879 | 28.4% |
| Media | 20,357 | 5,186 | 25.5% |
| Institutions | 7,703 | 957 | 12.4% |
| Reference | 15,643 | 1,171 | 7.5% |
| Forums | 23,986 | 1,268 | 5.3% |
Pages opened and pages cited by source type, 20,879 answers between 2 June and 22 August 2026.
No category gets past four in ten. Even social media, the best converter here, loses six of every ten pages the model opens.
At the bottom, forums and reference sites are doing a different job. The model uses them to work out what is true and then quotes somebody else.
The domains it reads most and cites least
Categories hide a lot. These are individual sites, all of them large and public, ranked by how little of their traffic survives into the answer.
| Domain | Type | Opened | Cited | Reaches the answer |
|---|---|---|---|---|
| en.wikipedia.org | Reference | 3,230 | 15 | 0.5% |
| es.wikipedia.org | Reference | 3,490 | 35 | 1.0% |
| cadenaser.com | Media | 923 | 14 | 1.5% |
| reddit.com | Forum | 23,433 | 1,060 | 4.5% |
| arxiv.org | Reference | 3,537 | 245 | 6.9% |
| elpais.com | Media | 2,062 | 182 | 8.8% |
| cincodias.elpais.com | Media | 2,067 | 195 | 9.4% |
| techradar.com | Media | 2,115 | 300 | 14.2% |
| semrush.com | Company | 1,446 | 355 | 24.6% |
| ahrefs.com | Company | 1,515 | 421 | 27.8% |
Large public domains only. Client sites and their direct competitors are excluded.
The English Wikipedia is opened 3,230 times and cited fifteen. That is not a source, it is a reference book. The model checks a fact, closes the tab and writes the sentence in its own words.
Media sit in an awkward middle. TechRadar converts at 14.2%, the Spanish national press at around 9%, and a radio brand at 1.5%. The model reads the news to understand a topic and then links to whoever sells the product.
The two best performers on the list are company sites. Ahrefs and Semrush are opened around 1,500 times each and cited on roughly one in four of those reads. Fewer reads, far more links.
Reddit is read more than anything else and shown less than almost anything
Reddit is the single most opened source in the dataset. 23,433 reads. It is also close to the bottom for citations: 1,060 links, 4.5%.
It is effectively the entire forum category on its own. Of the 23,986 forum pages the engines opened, 23,433 were Reddit.
Both facts are true at the same time, and they pull in opposite directions. What gets said about you on Reddit reaches the model. It rarely reaches the reader.
That changes what forum work is for. It is not a traffic channel. It is where the model forms an opinion about your category before it recommends anyone, and you will not see it in your analytics.
What to do with this
- Measure forums as reputation, not acquisition. Reddit shapes the answer and hands out almost no clicks. Judging it by referral traffic will tell you it does nothing, which is wrong.
- Keep your own pages as the main asset. Company sites convert at 33.2%, and the best individual performers in the whole dataset are two software vendors explaining their own products.
- Stop chasing reference sites. Wikipedia and arXiv are read constantly and cited almost never. Being there may help the model understand you. It will not send you anyone.
- Track what gets opened, not only what gets linked. Counting citations alone hides nearly three quarters of the pages involved in every answer, including the ones deciding whether you get recommended.
In one line
- Of 11,162 domains the AI opened, 6,686 (59.9%) were never cited once.
- Forums are 11.2% of what it reads and 2.7% of what it shows. Reddit converts at 4.5%.
- The English Wikipedia is opened 3,230 times and linked fifteen.
One caveat on the numbers. They come from 14 projects across 7 sectors, mostly software, accommodation and professional services in Spain. Treat them as a bearing, not a universal law. The split changes a good deal from one sector to the next, which is why the useful version of this exercise is the one run on your own data.
See what the AI reads about you and never shows
Bee LLM logs both layers every day: the pages ChatGPT, Gemini, Perplexity and AI Overviews link in your category, and the ones they only open. The free plan checks ChatGPT daily, with no card.
Start freeRead next: two in three sources the AI cites are company sites.
