215,128 Machine-Made Pages Are Grounding AI Answers: Inside the Trellner Study of Perplexity's Citation Supply Chain
Trellner Research ran 380 buyer-intent queries through Perplexity's sonar models and found 59.8% of 7,534 citations pointing to domains outside the world's top 100,000 sites — with three machine-generated 'best software' farms supplying 215,128 pages explicitly titled 'Facts & Grounding Page' for models to read.
On September 2, 2026, a research report landed with an unusually precise number in its title: 215,128. That is the count of machine-generated “best software” buying guides published by three websites that, between them, became a meaningful slice of the evidence base behind Perplexity’s product recommendations. The report — TR-2026-009, “Manufactured sources behind AI recommendations,” by Trellner Research — is the most detailed anatomy yet of a threat the AI industry has been discussing in the abstract for two years: search engine optimization aimed not at human eyes, but at the retrieval step of AI answer engines.
The findings deserve the same scrutiny the study itself applies to its subjects. But at face value, they describe a supply chain problem at the heart of AI search: when a model “grounds” its answer by fetching documents, the documents it fetches may have been manufactured for exactly that purpose.
What the study did
The methodology is refreshingly concrete. On September 2, 2026, Trellner put 380 buyer-intent software categories — ranging from “CRM software” to “museum collection management software” — to two of Perplexity’s models, sonar and sonar-pro, accessed through OpenRouter. One prompt per category per model, 760 calls in total. Each call asked for a ranked top-five list as JSON, with each product’s official homepage domain.
All 760 calls returned parseable answers, and both models report the URLs they retrieve — the property that made them suitable subjects. The categories were written before any results were seen and never revised, a pre-registration discipline rare in this genre.
The harvest: 3,800 recommendation slots naming 1,807 distinct products, and 7,534 citations spanning 2,055 distinct domains. Every cited domain was then looked up in the Tranco list (a standard ranking of the top million domains by real-world popularity) for September 1, 2026, and in the Wayback Machine. All 1,502 vendor homepages the models supplied were fetched — twice, once directly and once through a rotating proxy — to check whether they still exist.
Google was deliberately left out: grounding a Gemini model through OpenRouter routes through OpenRouter’s own web-search plugin, so the citations would describe that plugin rather than Google’s retrieval. The study’s scope discipline is worth emphasizing because it is what makes the numbers credible: only Perplexity was measured, and nothing in the report claims otherwise.
Where the citations land
The headline distribution is stark:
- 59.8% of the 7,534 citations point at domains ranked worse than #100,000 in Tranco.
- 23.4% point at domains not in the top million at all.
- The median Tranco rank of citations pointing at a ranked domain is 71,611.
Concentration at the top is unremarkable — the ten most-cited domains take 17.3% of citations, so the story is not a cartel of famous sites supplying the answers. The story is what fills the other four-fifths: 751 of the 2,055 cited domains — 36.5% — do not appear in the top million.
Those domains are also newer. The median first Wayback capture is 2020 for unranked cited domains versus 2011 for ranked ones, and 16.6% of the archived unranked domains were first captured in 2025 or later, against 1.6% of archived ranked domains. The long tail of AI grounding is a young web — one that grew up alongside the answer engines now citing it.
For comparison, Wikipedia was cited three times in 7,534 citations.
The vendor blog that became the third-largest source
The single most telling case study in the report is not one of the fake farms. It is Guideflow, a company that sells interactive product demos. It is not a review site, a directory, or a publisher, and it competes in none of the 380 categories. Yet its marketing blog was cited 194 times across 96 of the 380 categories — a quarter of them — placing it third overall, ahead of Gartner.
Each citation is a different URL: 96 distinct blog URLs, one per category, six of them the Estonian-locale copy of a post. Its sitemap lists 3,351 blog URLs, 2,176 of them distinct posts. It supplied the grounding for “3D rendering software,” “IVR software,” “RFID software,” and “architecture practice software” alike.
As the report is careful to state, nothing Guideflow does is deceptive — it publishes a large content-marketing blog, as thousands of companies do. The finding is about what the retrieval layer does with it: a vendor’s own listicles about markets it does not operate in became the third-largest evidence base for questions about which products to buy.
“Facts & Grounding Page”: content farms address the machine
Then there are the farms themselves. Three sites — Gitnux, Worldmetrics, and WifiTalents — were cited 71, 50, and 60 times respectively (181 citations total, 2.4% of the corpus), appearing in 41 of 380 categories.
The evidence that they are one operation is circumstantial but dense. All three were registered through NameCheap between December 2023 and May 2024. All three delegate DNS to the same pair of Cloudflare nameservers (pam.ns.cloudflare.com and sean.ns.cloudflare.com). All three run the same page template with the same navigation — Services, Market Data, Software Advice, Editorial Process, Company — and each keeps a blog of exactly six posts, all eighteen of which are about the other brands in the set. A fourth brand on the same nameserver pair carries the same homepage title.
Their scale is the point. Their sitemaps list 103,578, 107,083, and 105,541 URLs — of which 70,731, 71,684, and 72,713 follow the pattern /best/<something>-software/. That is 215,128 generated buying guides across three brands, against six blog posts each. As the report dryly notes: there are not 215,128 software categories.
What makes these sites unusual is their self-description. Fetched on September 2, 2026, two of them return an HTML title of the form <Brand> — Facts & Grounding Page, with an identical meta description apart from the brand name: “Verified facts about [Gitnux] … Company, legal, methodology, and compliance details in one machine-readable record.”
“Grounding” is not a term buyers use. It is the name of the step in which a retrieval system fetches documents to condition an answer on. A machine-readable record of verified facts about oneself is not a service to a human reader. These pages are addressed, in their titles and descriptions, to the software that reads them. And that reading is purchased in the ordinary way: one of the sites advertises custom market research “from €5,000,” ready-made reports “from €499,” and vendor selection “from €2,500” — sitting above the same taxonomy of generated Best Lists the models retrieve.
One template, three verdicts
To test the editorial substance, Trellner fetched the same category page — “project estimation software” — from all three brands. Each page states its ranking in JSON-LD, so it can be read without interpretation. Each ranks ten tools, with the top five displayed.
Gitnux’s winner does not appear in Worldmetrics’ top five at all. Each page carries three named staff — nine distinct people for one question. Each announces an editorial process; Gitnux labels its result “AI-verified · Expert reviewed.” All three carry an unrendered template variable in the byline line, reading “Within the next 26 days” on two of them and “Within the next 40 days” on the third — a telltale of mass generation that no human editor caught, because no human editor exists.
Where the recommendations themselves point
The 1,502 vendor homepages are mostly fine, and the study checked each twice. Ten resolve to no address at all — eight not delegated to any nameserver — including domains offered as the home of GraphiQL, Microsoft To Do, and Trivy. Another 92 (6.1%) redirect to a different registrable domain, mostly ordinary acquisitions and rebrands.
Two redirects are stranger, and in both cases the two model tiers disagreed. Asked for research data management platforms, both named Dryad — but one version of the cited domain redirected to an Indonesian online-gambling portal titled “BIGSLOT288 | Portal Game Online.” Asked for data quality tools, both named Monte Carlo — one citation redirecting to the Monte-Carlo Société des Bains de Mer, the Monaco hotel and casino group.
What the study does not show
The limitations section deserves as much attention as the findings, because it is what separates measurement from FUD:
- The two models are not two independent measurements. They returned byte-identical citation lists in 289 of 380 categories, with URL sets overlapping at a Jaccard similarity of 0.898. The Perplexity tiers share a retrieval layer and should be read as one search stack sampled twice.
- The result covers Perplexity only. ChatGPT, Gemini, Copilot, and Google’s AI Mode were not measured, and there is no basis for assuming their retrieval mixes match.
- The 380 categories are the researchers’ own construction, weighted toward niche verticals — which surfaces more long-tail sources than common queries would.
- Citation is not endorsement of correctness. The study did not test whether removing these sources would change the recommendations; the farms may name reasonable products. What was measured is which documents the evidence base is made of.
- Tranco rank is a popularity measure, not a quality measure. A low rank is not an accusation, and every claim about a specific site rests on that site’s own pages, linked and archived in the dataset.
The full dataset — every citation, every recommendation, the Tranco and Wayback lookups, the vendor liveness checks, and the scripts that produced every figure — is published under CC BY 4.0.
Analysis: GEO arrives, and the ground truth gets shaky
The term for this already exists: GEO, or generative engine optimization. Classic SEO games a ranking a human eventually sees; GEO games the retrieval step — the invisible moment where a model picks its sources. A content farm that could never reach Google’s first page can still end up cited, by name, inside a confident AI answer, because the retrieval layer is optimizing for “a page that matches this query exists,” not “a source a person would trust.”
The economics are brutal and asymmetric. A generated page costs fractions of a cent. A citation inside an AI answer is worth real money — the report notes companion coverage finding the median company appears in just 16% of the AI answers it would want to be in and is cited in 6%. The answer-engine page is becoming a manufactured format, and legitimate marketers are now competing with content farms that already know what the retriever likes.
There is a deeper structural problem. Answer engines create the demand for grounded citations, and the web responds by manufacturing supply. As AI answers become a dominant interface for product research, the incentive to build pages addressed to models — not humans — compounds. This is a reverse feedback loop: the more models ground, the more the ground truth is manufactured for them. Wikipedia, the closest thing the web has to a neutral knowledge commons, appeared three times in 7,534 citations — while two “Facts & Grounding Page” farms and a vendor blog outranked it.
For developers and enterprises building on retrieval-augmented systems, the practical takeaways are concrete. Domain-reputation priors (like Tranco rank) are cheap and surprisingly discriminating signals for retrieval filtering. Citation diversity requirements prevent any single source family from dominating. And the study’s methodology — pre-registered queries, full citation logging, independent liveness checks — is a template any RAG operator can steal to audit their own retrieval mix.
For everyone else, the advice is simpler. When an AI search engine gives you a sourced answer, the citation is doing a lot of quiet work to make it feel authoritative. Click through before you trust. The model checked that a page existed; it did not check that the page deserved to.
Sources
- [1] https://trellner.com/reports/manufactured-sources-behind-ai-recommendations/
- [2] https://onthewire.ai/article/215-128-fake-best-software-pages-were-built-to-be-read-by-ai-and-perplexity-is-c
- [3] https://artdirectiondaily.com/issues/2026-09-02-webflow-stops-requiring-webflow.html
- [4] https://www.explainx.ai/blog/ai-recommendation-sources-manufactured-geo-farms-trellner-2026