Product Thinking

Why AI Citations Are Not a Quality Signal

Bill Cava/

The third most-cited source Perplexity used to recommend software last month was guideflow.com, the marketing blog of a company that sells interactive product demos. It is not a review site, not a directory, not a publisher, and it competes in none of the 380 software categories tested.

It was cited 194 times, across 96 categories, one blog URL per category, six of them the Estonian-language copy of the same post. That placed it ahead of Gartner.[1]

Nobody at guideflow decided to become an authority on 3D rendering software. The retrieval layer decided for them.

That anecdote comes from one of two independent audits published the same day, September 2, by two different research shops, measuring opposite ends of the same pipe. Together they say something the growing get-cited-by-AI industry does not want to hear: a citation is a retrieval artifact, not an endorsement.

Do AI citations mean your content is trusted?

They mean it was retrieved. AI answer engines cite sources, and people have started counting those citations the way they once counted backlinks: as a proxy for authority. Counting is fine. Treating the count as a quality verdict is the error, because the layer producing it is measurably gameable at one end and measurably unfaithful at the other.

There is an entire discipline forming here (generative engine optimization, the SEO of AI answers), with services, software, and how-to-measure guides.

The demand is real, and the echo of the backlink era is hard to miss. But a proxy metric is only as good as the process behind it, and when the metric can be gamed, the score is theater. The two audits below are the first serious look at this particular process.

Where do Perplexity's citations actually come from?

Mostly from the web's long tail. Trellner Research asked Perplexity's two search models for the best products in 380 software categories and kept every URL: 7,534 citations across 2,055 domains. 59.8 percent pointed at domains ranked worse than #100,000 on the Tranco list of popular sites; 23.4 percent were not in the top million at all.[1]

The median popularity rank of the cited domains that had a rank at all: 71,611. Wikipedia was cited three times in 7,534. And it is not that a few famous sites dominate: the ten most-cited domains take just 17.3 percent between them. Here is that top ten.

Domain
Citations
Tranco rank
Where it sits
g2.com
291
4,027
Top 10,000 site
reddit.com
261
105
Top 10,000 site
gartner.com
158
1,766
Top 10,000 site
zapier.com
82
2,919
Top 10,000 site
capterra.com
68
6,387
Top 10,000 site
linkedin.com
67
18
Top 10,000 site
guideflow.com
194
177,039
Past rank 40,000
wifitalents.com
71
105,281
Past rank 40,000
worldmetrics.org
60
104,737
Past rank 40,000
gitnux.org
50
42,759
Past rank 40,000
Trellner's ten most-cited domains: citation counts in one band, popularity ranks four orders of magnitude apart.

Then there is the bottom of that table. Three of the top ten, wifitalents.com, worldmetrics.org, and gitnux.org, appear to be one operation: all registered through the same registrar between December 2023 and May 2024, all pointing at the same pair of nameservers, all running the same page template. Between them they host 215,128 machine-generated "best software" pages.[1]

Two of them give their homepage the same HTML title: "Facts & Grounding Page." Grounding is the technical name for the retrieval step these models perform. The sites are named for the machine they were built to be read by.

The worldmetrics.org homepage, a machine-generated statistics site whose HTML title reads Facts and Grounding Page
One of the three: 'proprietary data,' 'cited by hundreds of publications.' The page's HTML title, checked the day this post was written, is 'Facts & Grounding Page.'

One scope limit, stated plainly because the report states it plainly: only Perplexity was measured, and nothing here should be read as a claim about any other engine.[1] The lesson is not about one company. It is about what a citation count can and cannot tell you.

Do the cited pages even say what they're cited for?

Often, no. The same day, a second shop measured the other end of the pipe. Haus Research asked Perplexity's models 310 factual questions about 210 technology companies, fetched every cited URL, and checked whether the page contained any number from the sentence citing it. 34.7 percent of the 1,826 number-bearing citations failed.[2]

That figure is a floor, not a stretch. One matching figure anywhere on the page passes. A bare year passes. $185 million, $185M, and 185000000 all match. Blocked fetches got three attempts, the third through a rotating proxy that rescued 192 URLs, and the classification, in the authors' words, can only ever move in a page's favor.[2]

It is not link rot either. Just 1.3 percent of cited URLs were dead. The bigger category, 16.1 percent, sat behind logins, paywalls, and bot walls: pages a reader cannot open at all.

A footnote a reader cannot open is a claim of provenance with no way to test it, which is the condition a citation exists to prevent.

Haus Research, Perplexity citation audit, September 2, 2026

Put the two studies together and the pipe is polluted at both ends: the supply is partly manufactured for the machine, and the attachment between claim and source fails a third of the time it can be tested. Either study alone is a story about one engine having one problem. The pair is a story about the metric.

What survives if citations can be farmed?

Firsthand proof. In July we argued that expertise stopped being a moat once answer engines started delivering consensus knowledge wholesale, and that what is left is what you specifically did, measured, and can show. Back then the sharpest number available was that only about 38 percent of AI citations came from top-10 organic results, down from 76.

These audits supply what that post could not: a look at where the citations actually go. And the answer strengthens the case.

A template farm can produce 215,128 plausible category rankings in eighteen months. What it cannot produce is your deployment story, your measured outcome, your before-and-after. The same logic applies inside your own business: measure the thing, not the proxy, because proxies are exactly what gets manufactured first.

Should you stop doing generative engine optimization?

No. Being retrievable is worth engineering for, the same way being crawlable always was. What changes is how you read the result. Being retrieved is not evidence the work is good, and a metric that a page farm can move is not a measure of quality.

The practical version: if you report an AI-citation count internally, report next to it what the citing pages actually said. A third of the time, in the one audit that checked, the page did not say the thing.

And the uncomfortable version, applied to us. Our own AI-response count is the fastest-growing number on this site. We track it, we like watching it climb, and after this week it gets read the same way we are telling you to read yours: as a measure of retrieval. The verdict still has to come from somewhere else.

References

Frequently asked

How do you measure generative engine optimization?
Carefully, because the thing you are counting is partly manufactured.
Carefully, because the thing you are counting is partly manufactured. Two independent audits published on 2026-09-02 found that 59.8% of Perplexity's citations point at domains ranked worse than #100,000 on the Tranco list, and that 34.7% of citations attached to a sentence stating a figure led to a page that would not open or did not contain a single number from that sentence. A citation count measures retrieval, not endorsement.
Is generative engine optimization the same as SEO?
No, and the difference is not the one most guides describe. Classic SEO competed for a ranked list that a person then chose from.
No, and the difference is not the one most guides describe. Classic SEO competed for a ranked list that a person then chose from. Retrieval picks sources a reader usually never sees, which removes the check that made ranking mean something. Trellner found the median Tranco rank of Perplexity's ranked citations was 71,611, and Wikipedia was cited three times out of 7,534.
Does being cited by Perplexity or ChatGPT mean my content is trusted?
It means it was retrieved. In Trellner's audit the third most-cited domain was Guideflow, a vendor's marketing blog that sells interactive product demos, competes in none of the 380 categories tested, and was cited 194 times across 96 of them, placing it ahead of Gartner.
It means it was retrieved. In Trellner's audit the third most-cited domain was Guideflow, a vendor's marketing blog that sells interactive product demos, competes in none of the 380 categories tested, and was cited 194 times across 96 of them, placing it ahead of Gartner.
What actually survives if citations can be farmed?
Firsthand proof. A template farm can generate 215,128 pages of plausible category rankings, and three domains registered between December 2023 and May 2024 did exactly that.
Firsthand proof. A template farm can generate 215,128 pages of plausible category rankings, and three domains registered between December 2023 and May 2024 did exactly that. What it cannot generate is what you specifically did, measured, and can show.
Should I stop doing generative engine optimization?
No. Stop treating the citation count as the outcome.
No. Stop treating the citation count as the outcome. Being retrievable is worth engineering for; being retrieved is not evidence that the work is good, and a metric that a page farm can move is not a measure of quality.
Work with us

Let’s build it together.

We turn clever prototypes into production systems people can rely on. If you’re building with agents and want a hand making it real, leave your email and we’ll be in touch.

Straight to the team. No spam.