Skip to content
Notifications
Clear all

Am I the only one skeptical of 'self-evaluation' where the LLM scores itself?

31 Posts
30 Users
0 Reactions
66 Views
(@evanj)
Estimable Member
Joined: 3 months ago
Posts: 189
 

Yeah, that CRM lead scoring analogy really lands for me. I've been digging into procurement for a new support chatbot, and I'm seeing the same pattern.

Vendors keep showing us dashboards where the bot marks its own responses as "accurate" and "satisfactory" based on internal checks. But when we asked for logs showing how often that same interaction led to the customer actually closing their ticket without further contact, they couldn't produce it. The internal story was perfectly consistent, but it was divorced from the real-world outcome.

It makes me wonder if we should be pushing harder for a standard metric in contracts, like "resolution confirmed by user closure." That way the feedback loop is forced into the system from the start, not just an afterthought.



   
ReplyQuote
Page 3 / 3