We regularly publish aggregated, anonymised figures from our platform. Here, a simple question no one answers with real data: when an AI cites a source, does that link actually exist? And which model gets it wrong most often?
The scope
Between September 2025 and July 2026, the platform analysed 7,774 AI answers from 132 audits, produced by 53 distinct models (ChatGPT, Claude, Gemini, Perplexity, Grok, Mistral and others). Within these answers, the models cited 26,435 URLs as sources. We checked every one: 70 did not exist — plausible links to pages that were never published, spread across 23 different brands. An average rate of 0.26%, roughly 1 cited URL in 380.
The method is deliberately hard to dispute: a URL is counted as "invented" only if the model cited it as a source and it returns an error (404 or soft-404) at verification time. No interpretation: a page exists, or it does not.
The ranking — invented-URL rate per model
The only honest figure is not the raw error count, but the rate: what share of the URLs a model cites fails to resolve. Because a model that cites many links (like Perplexity) mechanically has more chances to be wrong. So here are invented URLs relative to total URLs cited, for each model with enough data.
| Model | Invented-URL rate | Invented URLs | Brands affected | URLs cited |
|---|---|---|---|---|
| Perplexity Sonar Pro | 1,05 % | 21 | 9 | 2 002 |
| Anthropic Claude Sonnet 4.6 | 0,62 % | 9 | 5 | 1 449 |
| OpenAI GPT-5 Mini | 0,44 % | 24 | 13 | 5 436 |
| Anthropic Claude Haiku 4.5 | 0,19 % | 8 | 3 | 4 198 |
| OpenAI GPT-5 | 0,16 % | 3 | 2 | 1 848 |
| Google Gemini 2.5 Flash | 0,12 % | 2 | 1 | 1 656 |
| Perplexity Sonar | 0,04 % | 3 | 3 | 7 065 |
How to read it: of the 2,002 URLs cited by Perplexity Sonar Pro in our audits, 21 (1.05%) did not resolve, across 9 brands. Models with 3 invented URLs or fewer are indicative — the sample is too small to rank finely, and we say so.
Finding 1 — The "search engine" paradox
The model that invents the most URLs, proportionally, is Perplexity Sonar Pro — precisely the one sold as a search engine that cites its sources. Yet the same vendor's lighter version (Perplexity Sonar) posts the best score of the whole panel (0.04%). The gap between the two is about 25×. The lesson: citation reliability doesn't depend on the model's brand, but on its precise configuration. That is exactly why measuring model by model matters — a generic "ChatGPT" or "Perplexity" means nothing.
Finding 2 — In absolute terms, the most-used model produces the most errors
GPT-5 Mini produces the highest raw count of invented URLs (24, across 13 brands) — not because it is the least reliable (its 0.44% rate sits mid-table), but because it is by far the most-queried model in our audits. In other words: at scale, even an average rate yields a substantial volume of errors. That is the trap of "default" models.
Finding 3 — No model is immune
Every family in our data — OpenAI, Anthropic, Google, Perplexity — produced at least one invented URL. There is no "clean" model at 0%. And a URL invented by a citation-first AI is more dangerous than a plain error: the user trusts the displayed source. In our 2026 barometer, we found 8 of these phantom addresses that received real human visits — prospects clicking and landing on an error page attributed to your brand. It is a reputation blind spot, detailed in our hallucinated URLs guide and the AI hallucination term.
What it changes for a brand
Two concrete takeaways. First, the model you are visible in matters as much as being visible: a citation in a model that invents links exposes you to erroneous associations. Second, the only way to know is to measure per model and verify every cited URL — which no manual monitoring can do at this scale.
The limits of these figures
This data comes from brands that chose to audit themselves — a sample that is not representative of all companies. Models with a low error volume (3 URLs or fewer) are indicative, not definitive. Rates evolve with every model update; we will republish this study with fresh data. The full methodology, including what we do not measure, is on our methodology page.
Want to know which models cite — or invent — URLs for your brand? That is exactly what an audit measures, model by model, URL by URL.
Every question asked to ChatGPT without your name in the answer is a competitor recommended instead of you — measured across 6,820 real AI answers.