You've put your finger on the exact issue: the feature pipeline *is* the model. The rounding to city center isn't a bug, it's a design choice that abstracts away the very signal you need.
I've seen this happen with timestamps, too. A platform rounded all login times to the nearest 15-minute block to reduce cardinality for its clustering algorithm. It completely missed rapid, sequential logins from different tills during shift changes because the events were bucketed together. Their "impossible travel" between stores was just a cashier logging out and their relief logging in.
>their ML is built on a shaky foundation
It's often worse. Sometimes the foundation is solid, but they've built a house on stilts next to it. The raw data is fine, but the feature transformation layer is a proprietary black box they're afraid to let you touch because it's their "secret sauce." In reality, the sauce is just a bunch of arbitrary filters that don't fit your business context.
That 15-minute timestamp rounding is a perfect example of data loss masquerading as a feature. It's not about cardinality, it's about lazy engineering.
It kills detection for any fast-paced retail process. Think point-of-sale refund loops or inventory system scans. You can't detect rapid-fire malicious activity if your clock resolution is a quarter hour.
>The raw data is fine, but the feature transformation layer is a proprietary black box
That's the real lock-in. You can't fix their bad math. You're stuck with their broken model, paying for it every month, while your team builds workarounds outside the platform. At that point, you're just buying a fancy data collector.