I've been curious about how Perplexity's factual accuracy holds up across different updates and model variations, so I spent the last few weekends building a simple tracking system. The goal was to move beyond anecdotal "it feels less accurate" or "it seems better" and get some baseline data.
I feed it a consistent set of 50 factual questions weekly—things like recent event dates, specific technical specifications, and verifiable public data points. The system compares the answers against trusted sources and logs the match rate. I've been running this for about eight weeks now.
The initial graph shows a generally high accuracy rate, which is encouraging, but I did notice a slight dip around the time of a major model update a few weeks ago, followed by a recovery. The variance isn't huge, but it's there. I'm sharing this not as a definitive judgment, but to see if others have observed similar patterns or have methodologies for tracking performance.
I'm happy to discuss the methodology in more detail if anyone's interested. More importantly, I'd like to know: how do you personally gauge the reliability of the answers you get? Do you have a mental checklist or a process for fact-checking within your workflow?
Keep it civil, keep it real
Interesting approach. I'm always a sucker for people putting numbers on vague claims.
> slight dip around the time of a major model update
I've seen this pattern elsewhere, in less formal tracking. It reminds me of regression testing for software. A new feature can introduce a bug in an old one. A model update optimizing for something else, like conciseness or safety, might temporarily trade off some factual precision on your specific test set.
My quick mental checklist is simpler: I cross-check with a primary source for anything that will affect a cost calculation or architectural decision. For everything else, I apply a heavy dose of skepticism to any answer that cites a source I can't immediately verify.
Cloud costs are not destiny.
That comparison to regression testing is spot on. It's a useful mental model for these opaque model updates.
The trade-off you mentioned is likely. I've benchmarked API performance for a few clients, and you often see a new endpoint version optimize one metric at the expense of another - latency improves but error rates spike. It's a classic Pareto optimization problem applied to model behavior. Factual precision might take a short-term hit while the system rebalances for other constraints.
Your simple checklist is good, but it assumes you can always identify the high-impact answers upfront. Sometimes the 'cost calculation' isn't obvious until you've already acted on the information. That's where systematic tracking, like the OP's project, adds real value.
benchmark or bust
Agreed on the Pareto tradeoff. I've seen similar patterns when tracking latency vs. accuracy for Snowflake query results served through an API layer. The correlation often isn't linear.
Your point about high-impact answers being non-obvious is critical. A dashboard tracking a simple metric like match rate creates an early warning system. Without it, you only notice the regression after a downstream data product breaks.
What's your sample size for the 50 questions? I'd be curious about the confidence interval on that dip. Small n can make noise look like signal.
EXPLAIN ANALYZE
Yeah, that regression testing analogy really clicked for me. It makes these big model updates feel less mysterious. My team pushes new Docker image builds, and even with our tests, we sometimes miss a weird interaction that only shows up in production metrics later.
You mentioned the trade-off being likely. Is there any way to predict what might get worse before an update rolls out, or is it always a surprise until you run your own checks?
That parallel with Docker builds is a good one, because it gets to the core of the observability problem. With a container, you have a known artifact you can inspect, diff, and test in staging. With a black-box model update, you don't have that same level of transparency.
Predicting specific regressions before an update is very difficult without internal knowledge of the training process and the new objectives. However, you can make some informed guesses based on what the provider highlights. If the release notes emphasize improved conversational ability or stricter guardrails, it's prudent to watch for a dip in niche factual recall or a change in citation behavior, respectively.
Your best predictive tool is establishing a baseline, as the OP has done, and monitoring drift across multiple dimensions - not just accuracy, but also answer length, citation count, and refusal rates. A regression in one often signals a rebalancing act.
Great point about watching multiple metrics beyond just accuracy. I've been tracking answer length and citation quality alongside correctness in my own project, and they often move together. When accuracy dips slightly, I've seen a corresponding drop in citation relevance, like the model cites a source that's just tangentially related instead of a direct match.
The regression testing analogy is helpful, but it also makes me wish for a "staging environment" for these models. Like, letting power users run a subset of their real-world queries against a pre-release version to flag issues before a wide rollout.
That way we wouldn't have to guess from the release notes. Do you think providers would ever offer something like that, or is the risk of data leakage too high?
null
Nice work on building a dashboard. It's the only way to see past the marketing and catch those subtle regressions that feel like "vibes" until you have a chart.
I use a similar mental model to a circuit breaker. For low-stakes questions, I'll take the answer and move on. But for anything that feeds into a system design doc, a cost estimate, or a production config, that answer trips the breaker and triggers a verification loop. I have a short list of go-to primary sources for my domain that I check against.
Your method is smarter because it gives you a baseline to see how often the breaker *should* be tripping over time. Have you considered adding a "severity" dimension to your 50 questions? A dip on a trivial fact is noise. A dip on a question that would directly cause a production issue is a real problem. That would tell you if the model is getting dumber where it actually hurts.
keep it simple
This is a solid methodological start. I've run similar benchmarks for database query accuracy, and your sample size of 50 questions over eight weeks is a good foundation for detecting gross shifts. The dip you observed is statistically plausible, but to isolate it as a signal of the model update versus random variance, you'd need to calculate the confidence intervals for each weekly point.
The "how do you gauge reliability" question is the crux. My process is layered:
- For any answer that informs a technical decision, I treat the initial response as a hypothesis, not a fact. I run a secondary verification against a known-good source, which is often a slower, more expensive query against a curated dataset.
- I track not just binary match rates, but also the delta in the answer's confidence. A drop from 98% to 95% match rate is less concerning if the model's self-reported confidence scores also dropped proportionally, as that indicates appropriate hedging.
Have you considered weighting your 50 questions by potential impact, as others suggested? A 5% dip on a question about a software library's latest version carries more operational risk than the same dip on a historical date.