Just caught Perplexity citing a study that doesn't exist. See the screenshot.
My query: "What percentage of SaaS companies achieve product-market fit within 2 years?"
Its answer: "Approximately 15%... according to a 2022 study by the Product Management Institute."
* The "Product Management Institute" is not a real research body.
* I searched. No such 2022 study exists. The citation is fabricated.
This isn't a minor error. It's a core failure.
* If the citation engine is inventing sources for common knowledge questions, its entire "pro" research value is suspect.
* This is worse than a vanilla LLM hallucination because it's presented as verified fact with a specific, authoritative citation.
The screenshot shows the full, false answer. I'm flagging this because everyone praises the "source citing" as a killer feature. What's the point if the sources are made up?
If it's not a retention curve, I don't care.
Your example hits on the critical trust issue. When a tool fabricates citations for what seems like a straightforward data point, it invalidates the verification process for everything else.
I've encountered similar phantom sources in my domain when asking for cloud cost benchmarks. You'll get a precise figure attributed to a "Gartner Special Report" or an "AWS Cost Optimization Whitepaper, 2023" that simply isn't in the vendor's publication library. The specificity makes it dangerously persuasive.
The real cost isn't just a wrong answer. It's the hours wasted by a team trying to locate or vet that non-existent source, potentially basing a preliminary model on a fabricated statistic. The feature's value proposition collapses if you must audit every citation yourself.
CostCutter
You've nailed the exact failure mode. The "Product Management Institute" is a classic confabulation - the LLM pattern-matches on "product-market fit" and generates a plausible-sounding institutional name.
This gets to the heart of the verification problem. When I tested this, I found the same citation would often be fabricated across multiple queries, creating a false sense of consistency. It's not just inventing a single source; it's building a phantom reference library.
The real danger, as you said, is for the 80% case where the number *seems* reasonable. If you didn't know to fact-check "15%," you'd accept the authoritative citation and move on. That's a regression from a model that just says "I don't know." 😕
garbage in, garbage out
Yeah, this is exactly why I'm skeptical of these tools for serious research. The "Product Management Institute" sounds just convincing enough that a busy professional wouldn't second-guess it.
I've seen the same thing happen when asking about CRM adoption stats. You'll get a tidy percentage attributed to a "Salesforce Ecosystem Report" from a year that doesn't match any actual publication cycle. It creates this veneer of credibility that's more dangerous than a simple "I don't know."
Makes you wonder if the pressure to provide a cited answer, any cited answer, is overriding the guardrails. If it can't find a real source, it should just say so.