← Back to all posts
News

Three Sites Wrote 215,128 Buying Guides for Robots. Perplexity Cites Them.

September 3, 2026 · 03:11 UTC · News
Three Sites Wrote 215,128 Buying Guides for Robots. Perplexity Cites Them.

TL;DR

Trellner Research asked Perplexity the same question 380 times, "What are the best {category} in 2026?", once each against its sonar and sonar-pro search models, then traced where the 7,534 citations behind the answers actually went. The report, published September 2, finds that 59.8% of citations point to domains outside the 100,000 most-visited sites on the web, that Wikipedia earned three citations in the entire run, and that three connected domains which have titled their homepages "Facts & Grounding Page" host 215,128 machine-generated buying guides between them. The full dataset and the scripts are public under CC BY 4.0.


What Trellner actually ran

The setup is small enough to reproduce over lunch. Trellner fixed 380 software categories in advance (B2B tools, developer infrastructure, vertical industry apps), fixed one prompt, and asked for exactly five ranked products back as JSON. Every call went through OpenRouter to the two Perplexity models, 760 calls in total, and every one of them parsed. That produced 3,800 recommendation slots naming 1,807 distinct products, backed by 7,534 citations spread across 2,055 domains.

Each cited domain was then looked up against the Tranco top-1M popularity list generated the day before, checked for its first capture in the Wayback Machine, and, for the 1,502 vendor homepages Perplexity supplied alongside its picks, tested for liveness and redirects. Trellner says it takes no money from the companies it covers, and it published every CSV, the raw answers, and the 13 Python scripts that produced each figure.

Who Perplexity cites for "best software"

The head of the distribution looks sane. G2 leads with 291 citations, Reddit follows with 261, and Gartner sits fourth with 158. Then it gets odd. Third place, ahead of Gartner, belongs to Guideflow, a company that sells interactive product demos, with 194 citations. And three names you have never typed into a browser sit inside the top ten: wifitalents.com with 71, worldmetrics.org with 60, and gitnux.org with 50. Wikipedia got 3.

citations per domain, 380 queries (7,534 total) g2.com291 reddit.com261 guideflow.com194 gartner.com158 wifitalents.com71 worldmetrics.org60 gitnux.org50 wikipedia.org3 copper = three linked template sites; slate = a vendor blog
A demo-software vendor's blog outranks Gartner, and three template farms outrank Wikipedia by a factor of 60.

Sixty percent from the long tail

Ranking every cited domain by popularity gives the headline number. 59.8% of citations point to domains ranked worse than 100,000 on Tranco, and 23.4% point to domains that do not appear in the top million at all. The median ranked domain sits at position 71,611. The top ten domains together account for only 17.3% of citations, so this is not a story about one bad source. It is a story about the evidence base being mostly places nobody visits.

Age tells the same story. Among cited domains that are unranked, the median first Wayback capture is 2020, versus 2011 for ranked ones, and 16.6% of the unranked, archived domains were first captured in 2025 or later. For ranked domains that figure is 1.6%.

share of 7,534 citations by Tranco rank of cited domain 40.2% top 100k 36.4% 100k-1M 23.4% past 1M 59.8% beyond the top 100,000 sites
Three in five citations come from the part of the web that popularity lists barely register.

Three sites, one template, 215,128 pages

The report's centerpiece is the trio inside the top ten. wifitalents.com, worldmetrics.org, and gitnux.org were all registered between December 2023 and May 2024, share the same two Cloudflare nameservers, use an identical page template with the same navigation, and each maintains exactly six blog posts about the other two brands. A fourth domain, zipdo.co, sits on the same nameservers. Counting their sitemaps, Trellner found 72,713, 71,684, and 70,731 pages under a /best/ path respectively: 215,128 "best X software" guides between them, on sites that did not exist before December 2023.

The detail that gives the report its title is the HTML title tag. Two of the three sites call their homepage a "Facts & Grounding Page", and the meta description introduces "an independent market research company publishing industry statistics." (We checked on September 3: the Worldmetrics and Gitnux homepages still carry that title.) As Trellner puts it, "Grounding is not a term buyers use. It is the name of the step in which a retrieval system fetches documents to condition an answer on." It is the first website we have seen that names its target audience in the title tag, and the audience is not you.

Grounding is the mechanism that makes this work. A search model like Sonar runs a retrieval step, pulls a handful of documents off the web, writes its answer from them, and attaches those documents as citations. Picture a hotel concierge who recommends restaurants from whichever flyers were pushed under the door that morning: the recommendation arrives with a source, but the source is the flyer, and whoever prints the most flyers gets recommended. Between them the three farms collected 181 citations, 2.4% of the total, and appeared as sources in 41 of the 380 categories.

best X in 2026?380 categories Sonar retrievalfetches grounding 215,128 /best/template pages 5 rankedproducts
The farms are built for the middle box: they exist to be fetched, not read.

Their homepages are not shy about the scale, either. Worldmetrics advertises 72,716 reviews across 96 categories, plus custom research from 5,200 euros. Gitnux says its work has been cited by Microsoft, Adobe, Google, and Harvard Business Review. WifiTalents claims the same of The New York Times, the Washington Post, and the Wall Street Journal. None of the three names an owner; Trellner infers common control from the shared infrastructure and the identical template.

/best/ pages per domain (sitemap count) wifitalents.com72,713 worldmetrics.org71,684 gitnux.org70,731 same template, same 2 nameservers, registered Dec 2023 - May 2024
Three brands, one production line: each domain carries almost exactly the same number of buying guides.

This is the second time in a month the grounding step has been caught eating manufactured input. In August, Politico and Responsible Statecraft exposed the Hanover Institute, a think tank that exists only as a website, publishing 100-plus AI-written reports built to be cited by chatbots. That operation was aimed at foreign policy. This one is aimed at your software budget.

The vendor blog that outranks Gartner

Guideflow is the more instructive case for anyone who ships software, because it is not a spam farm. It is a real company whose blog sitemap lists 3,351 URLs and 2,176 distinct posts, and that blog supplied grounding in 96 of the 380 categories, a quarter of the run, including categories with nothing to do with demos, like "3D rendering software" and "RFID software". A large enough pile of category pages on a legitimate domain makes you a source for the whole market, not just your own corner of it. That is a distribution lesson, whichever side of it you find distasteful.

Sonar and Sonar Pro are the same witness

If you were planning to cross-check one Perplexity model against the other, save the tokens. In 289 of the 380 categories (76%) the two models returned byte-identical citation lists, the Jaccard overlap across the whole run is 0.898, and they name the same top product in 290 categories. Trellner's wording is precise: the two models "are not two independent measurements" and "should be read as one search stack sampled twice." On OpenRouter, sonar-pro lists at $3 per million input tokens and $15 per million output, plus $5 per thousand searches, and that premium did not buy a different evidence base in this run.

Dead links and a casino

The liveness check on the 1,502 vendor homepages Perplexity handed back is the report's comic relief, unless you are one of the vendors. 17 domains (1.1%) were unreachable or do not exist, and 92 (6.1%) redirect to a different registrable domain. Some of those are ordinary brand moves. Some are not: dryad.co, supplied as a product homepage, now lands on "BIGSLOT288 | Portal Game Online", an Indonesian gambling portal, and montecarlo.com resolves to the Monaco hotel and casino group. Perplexity will cite a source for a product whose homepage now invites you to try your luck.

What the report does not prove

  • Perplexity only. ChatGPT, Gemini, Copilot, and Google's AI Mode were not measured, and Google Search was excluded entirely.
  • One run, one prompt. Each category was asked once, with one wording the authors wrote, through a datacenter proxy under a research user-agent. Real buyer phrasing and repeat sampling might move the lists.
  • Citations, not outcomes. Trellner's own line: "We have not shown that any of this changes the answers. What we measured is which documents the evidence base is made of."
  • Popularity is not quality. A Tranco rank past 100,000 says a site is obscure, not that it is wrong.
  • No response from Perplexity appears in the report, and we found none elsewhere at publication time.
  • The endpoint is moving. Perplexity's docs say the Sonar chat-completions API is being retired in favor of an Agent API, with support through September 27, 2026. Whether the retrieval layer behind the citations changes with it is not documented.

If you sell software

  • Find out what is grounding your category. The dataset has a citations.csv with every source per category. Filter it to your market and see whether Perplexity is reading G2 and Reddit or a template page generated last spring.
  • Search the farms for your own name. If you are in one of those 215,128 pages, you are being described by text no human wrote and nobody at the site will correct.
  • Own the category pages on a domain you control. Guideflow became the third-largest source in the study with a blog. A boring, comprehensive comparison page on your own domain is retrieval bait in the good sense.
  • Audit your own homepage. 6.1% of recommended vendor domains redirect somewhere else, and Perplexity keeps recommending them anyway. Do not be the casino.

Key Takeaways

  • Trellner ran 380 "best software" queries against Perplexity's sonar and sonar-pro and traced 7,534 citations across 2,055 domains, with the data and scripts public under CC BY 4.0.
  • 59.8% of citations point to domains outside the Tranco top 100,000, 23.4% are outside the top million, and Wikipedia was cited three times.
  • Three connected domains registered since December 2023 host 215,128 template "best X software" pages, and two of them title their homepages "Facts & Grounding Page".
  • Guideflow's product blog was the third most-cited source, ahead of Gartner, appearing in a quarter of all categories.
  • sonar and sonar-pro returned byte-identical citation lists in 76% of categories; treat them as one retrieval stack, not two opinions.
  • 6.1% of recommended vendor homepages redirect elsewhere, including one to an Indonesian gambling portal.

Sources: Trellner Research, TR-2026-009, Trellner dataset and METHOD.md, Perplexity model docs, OpenRouter sonar-pro, Tranco, Dr. Web

AIPerplexityAI SearchSEOGEOContent FarmsSaaSTrends
CONSOLE
$