Citation Farms: 215,128 Fake Pages Are Inside Your AI's Answers

October 1, 2026 · 6 min read

Citation Farms: 215,128 Fake Pages Are Inside Your AI's Answers

TL;DR - One operator registered three websites in late 2023 and filled them with 215,128 machine-generated "best software" buying guides. No human reads them. AI answer engines do: in a 7,534-citation study of what Perplexity retrieves, those three sites took 2.4% of all citations, and 59.8% of everything cited came from domains ranked worse than 100,000th on the web. Every check in an AI answer pipeline tests something that is now cheap to fake. Here is how the farm works, why Google's spam filters never saw it coming, and the four-move defence.

By The Numbers

NumberContext
215,128machine-generated "best software" pages across three sites (Trellner TR-2026-009)
3sites (gitnux.org, worldmetrics.org, wifitalents.com), all NameCheap-registered Dec 2023 to May 2024
2.4%share of all 7,534 citations in the study taken by those three sites
59.8%of all citations pointing at domains ranked worse than Tranco #100,000
3times Wikipedia appeared in 7,534 citations
~1 in 2answers that cited Ahrefs' test pages but recommended a competitor anyway
1 in 3eligible days an Ahrefs test page actually held its citation
1%of visits where a user clicks a source inside an AI summary (Pew, 900 US adults)
~16%of answer-engine citations already assessed as AI-generated (Northwestern audit)

Ask an AI assistant for the best CRM and you get a tidy answer with citations. Those citations carry a promise: somebody, somewhere, evaluated the options and wrote down what they found. A new kind of operation exists to break that promise at industrial scale, and it is not run by kids in a basement. It is run by a company with a VAT registration.

Inside the farm

In September 2026 the research firm Trellner asked two of Perplexity's retrieval models for the best products in 380 software categories and kept every source they pulled. Two of the sites doing the grounding, plus a third under apparently common control, turned out to be one operation: 215,128 pages of "best X software" buying guides spanning categories so combinatorial that "pick edit software" and "pick editor software" are separate guides. None of the three domains existed before December 2023. All three were registered through NameCheap within months of each other and share the same DNS.

The sites did not hide. One of them titled its homepage "Facts & Grounding Page". Grounding is retrieval jargon no software buyer has ever used, but it reads as reassurance to a machine counting sources. The operator behind all three is Global Commerce Media GmbH, a German company; the video that broke the story down found the parent organisation in the sites' own structured data. Eight brands, one operator. An engine counting citations saw independent consensus. It was one template talking to itself.

The business model is the quiet part. The same operation sells software advisory from 2,500 euros, with vendor shortlists drawn from the generated lists. One paying client covers the spam budget for years.

Every check is a stand-in

When an answer engine builds a response it runs checks at every stage. Look at what each one actually measures:

  • The crawler checks that a page is reachable. 215,000 pages pass.
  • Retrieval checks that words match your question. 215,000 exact-title pages are 215,000 precomputed matches for the smaller queries your question gets split into.
  • Freshness wants a recent timestamp. The farm regenerates dates daily, for free.
  • Quality scoring looks for an author bio and an editorial-process page. Generated headshots satisfy it.
  • Synthesis counts its sources instead of authenticating them. Three sites voting identically look like three independent opinions.

Every stage checks something measurable. None of them measures truth. Every one of those stand-ins is now cheap to fake, and a few hundred dollars of tokens buys the whole farm.

Why Google's spam filters never saw it

Google beat the content farms once. The 2011 Panda update and the anti-spam stack around it leaned on behaviour: how long people stayed, whether they bounced back to the results. A citation farm has no human visitors, so there is no behaviour to demote. The site is invisible to the strongest defence Google has, precisely because it was never built for humans.

There is a second gap. Google's demotions live in Google's index. An answer engine grounding on its own crawl inherits none of that fifteen-year cleanup. It starts the spam war from scratch.

The economics changed too. Classic spam competed for one of ten result slots against real competition. A farm needs one of five slots for "best museum collection management software", a query no honest publisher ever bothered writing. Nearly 60% of everything Perplexity cited sat outside the top 100,000 websites. Retrieval shops in the long tail, and the farms moved in next door.

Cited is not recommended

Here is where it gets stranger. Ahrefs ran the controlled version: 34 self-promotional pages across 5 domains, tracked through nearly 10,000 AI answers. In about half the answers that cited one of their pages, the model never mentioned their brand. It read the page as research and recommended a competitor straight off their own list.

The trick only fills a vacuum. A brand-new conference filled 72 empty answer slots, while an established product got 94% of its new mentions from other people's content. And a citation is a rental: Ahrefs' pages held theirs on only about one day in three. Even the spam does not reliably win.

Nobody clicks anyway

Pew Research tracked the real browsing behaviour of 900 US adults through March 2025. When a Google search produced an AI summary, users clicked a source inside that summary on 1% of visits. The citation is a trust badge almost nobody checks - which is exactly why planting one is worth doing.

A broader Northwestern audit of four answer engines put roughly 16% of cited sources as AI-generated across 712 real queries on politics, health and the environment. Your users are already consuming synthesized ground and calling it research.

The study that exposed it looked fake too

Hacker News roasted the Trellner report within a day of publication - the research firm's domain was freshly registered, the leadership team had AI-polished portraits and no footprint anywhere else, and the submitter's account had pushed four similar "research institutes" in one day. A story about manufactured sources arrived wearing the manufactured-source costume.

What saved it was the one thing a letterhead cannot fake: the full dataset and scripts shipped with the report, and the numbers reproduce when you rerun them. That is the bar now. The about page is the cheapest thing on the internet. The rerun is not.

What to actually do

  • Be the 1%. Click the citation before you act on the answer. Thirty seconds of source-checking beats a fabricated recommendation.
  • Prefer primary artifacts. Vendor documentation, changelogs and package registries sit one fakeable step closer to the truth than pages that describe them.
  • If you run retrieval (a RAG pipeline, an agent with web search), deduplicate your sources by operator, not by domain - one DNS lookup exposed this entire network - and require cross-operator agreement before a claim drives an action. Log what you retrieve so you can replay which source moved which answer.
  • Trust what you can rerun. A vendor's about page, their research report, their case studies: all free to manufacture. A dataset that reruns is not.

The farm spent a few hundred dollars. The defence costs attention, once, and then it is a habit.


Sources: Trellner Research TR-2026-009 "Manufactured sources behind AI recommendations" (Sep 2026); Ahrefs, "Self-promotional content works until it backfires"; Pew Research Center (22 Jul 2025); Allaham & Diakopoulos, arXiv:2605.23684; Devsplainers, "How To Spam Your Way Into AI Recommendations (Citation Farms)".

Mathew Clark Founder, SecureInSeconds Currently: deduplicating my reading list by operator.

Share:
Buy me a coffee

You might also like