Yeah, that CRM lead scoring analogy really lands for me. I've been digging into procurement for a new support chatbot, and I'm seeing the same pattern.
Vendors keep showing us dashboards where the bot marks its own responses as "accurate" and "satisfactory" based on internal checks. But when we asked for logs showing how often that same interaction led to the customer actually closing their ticket without further contact, they couldn't produce it. The internal story was perfectly consistent, but it was divorced from the real-world outcome.
It makes me wonder if we should be pushing harder for a standard metric in contracts, like "resolution confirmed by user closure." That way the feedback loop is forced into the system from the start, not just an afterthought.