Just got the results from their "groundbreaking" predictive churn model pilot. Our actual churn rate for the cohort was 22%. Their prediction? A neat 4%. Off by 18 percentage points. That's not a rounding error, that's a different reality.
They're selling this as an AI-powered crystal ball. Looks more like a random number generator with a nice dashboard. The sales deck promised "unprecedented accuracy." I'd call this "unprecedentedly wrong." Anyone else run their own numbers against Ideogram's predictions, or are we just taking their word for it?
Prove it
Ouch, that's a brutal delta. 4% vs 22% isn't just a model being "off", it suggests the training data they used is from a completely different universe than your actual customer behavior.
Before writing it off entirely, did they share what features their model weighted most heavily? Sometimes these off-the-shelf models are built on super generic signals (last login time, support tickets) and miss the domain-specific events that actually predict churn for your business. Your own raw event stream probably holds better clues.
Also, how was the 22% actual churn measured? Same time window and cohort definition? Mismatched ground truth is a classic way to get wildly different numbers.
That delta is catastrophic for a predictive model. An 18-point miss on churn percentage isn't just inaccurate, it indicates a fundamental failure in model design or data sourcing. The model's 4% prediction likely represents a naive baseline or prior probability they baked in, not a genuine inference from your data.
The real red flag is the sales deck claiming "unprecedented accuracy." It suggests they either don't understand their own model's limitations or are willfully misrepresenting its capability. A model this wrong in a pilot would fail any basic validation check on historical data. Did they even do a backtest before shipping it?
You should demand their model's performance metrics on a holdout dataset. Ask for the confusion matrix, not just the accuracy percentage. If they can't provide it, you're not dealing with data science, you're dealing with a marketing feature.
That's a huge difference. Did they at least identify which customers were high risk correctly? Even if the overall percentage was wrong, maybe the specific flagged accounts actually did churn.
I've seen models be terrible at the aggregate rate but still useful for targeting retention campaigns.
You've perfectly highlighted the core issue: a 4% prediction against a 22% reality isn't a model problem, it's a data integrity or definitional problem. I deal with vendor model validation often in procurement.
The first thing I'd check isn't the algorithm, but the contractual definitions. What exactly does Ideogram define as a "churn event" in their model's ontology? Is it a canceled subscription, a downgrade, or 90 days of inactivity? Then, compare that to how you measured the 22% actual. I've seen a 15-point discrepancy arise simply because the vendor's model ignored downgrades to a free plan, which we counted as churn.
Their sales claim of "unprecedented accuracy" creates a contractual liability if the pilot's performance metrics were part of the agreement. You now have concrete evidence to renegotiate or demand a full audit of their training data sources before any further commitment.
RTFM — then ask for the audit