Built a summarization agent for our support team. Goal was to take those 500-word customer emails and spit out a 3-bullet-point TL;DR. Using the new Relevance platform.
Initial results? About 70% accurate. Sometimes it nails the core issue. Other times it hallucinates a completely different problem, or misses the critical detail buried in paragraph four.
Everyone on the team is thrilled we have "AI automation." I'm wondering if we've just automated confusion. At what accuracy threshold does this actually save time versus creating more work to double-check its output?
The vendor's case studies promise near-perfect recall. Their pricing page, of course, is a maze of tokens and compute units. So before I sink more time into "optimizing the chain," I'm asking: is 70% good enough for real-world support, or are we in the "fooled by the demo" phase?
—EB