Yep. You're spot on about comparing to a known baseline. Without that, it's just a vanity number. I've seen this exact move in CRM analytics - a vendor claims "30% better lead scoring" but won't say if it's against a random baseline or Salesforce's native score. It makes the number useless.
The methodology is the product. If they can't share the test set, the claim is functionally worthless for any serious evaluation.
Right? It's exactly like when a CRM add-on claims "30% more qualified leads" but won't disclose the scoring rubric or the control group. The absence of a test set and a baseline is the loudest data point they're giving.
In my last integration project, I made it a rule to ask for the benchmark dataset upfront in the first sales call. If they couldn't provide it or a detailed methodology doc, the call ended there. It saved so much time.
That missing baseline - "compared to what?" - turns their biggest selling point into a giant question mark. You end up having to build your own evaluation, which defeats the purpose of buying a solution.
Exactly. No baseline, no metric, no test set. It's a null claim.
I treat it like a vendor claiming superior uptime but refusing to share their SLO calculation or incident logs. You can't integrate that into a real plan. It's a decorative statistic, not an operational one.
If they can't define "accuracy" in a reproducible way, they haven't measured it. They're just describing a feeling.