Spot on about the pitch being the real problem. When vendors frame a new tool as a *replacement*, it forces an unnecessary either/or decision for teams.
The power tool vs. screwdriver analogy is perfect. It reminds me of pitching a research phase. You don't use the same tools for discovery that you use for production. Trying to force the "extreme context" model into a stable workflow is like using a sledgehammer to tighten a screw - it might work, but it's inefficient and risks breaking what's already solid.
Maybe the pushback should be, "Show me how this fits in our *existing* process for exploratory work, not how it overthrows the proven one."
The research vs production tool split is exactly where we've seen some success. We used a large context model for analyzing a year's worth of unstructured support chat logs, something our standard classifiers couldn't touch.
It was strictly a one-off discovery project. We never connected it to a live API. The output just gave us a list of potential new categories to hardcode into our existing, simpler tagging system.
The key was keeping it completely separate from the operational stack. Once leadership saw the cost of running that model continuously versus the value of the handful of new rules we extracted, the push for a full integration died.
Latency is the enemy, but consistency is the goal.
You're right about the mismatch. I've been looking at help desk integrations, and I see the same pattern.
Vendors pitch these models for "understanding complex ticket sentiment," but our existing ticket tagging system, built on a simpler model, already routes 95% of issues correctly. The extreme context feels aimed at a theoretical problem where a customer submits a novel every time they need help.
Is anyone actually getting value from the ambiguous task handling, or is it mostly for pre-sales demos?
Exactly. We're onboarding a new team in Jira and I get pitched an AI add-on for 'context-aware' ticket routing almost weekly. But as you said, our current system works fine for nearly everything.
I've heard a few people say the ambiguous handling helped with weird, one-off escalations from long-term clients where the history matters. But that feels like it's solving for the 1% exception, not the daily workflow.
Do you think there's a way to trial these models just for that edge case without them taking over the whole tagging system?
Trial it for the 1%? That's the setup. They'll argue you can't prove value without seeing it handle the full volume, and then you're on the hook for the whole system.
If you're desperate to test, don't route tickets. Manually dump the weird ones into the model once a week and read the output yourself. If it's garbage, you've lost an hour. If it's brilliant, you've proven a manual review step works, not an automated takeover.
But honestly, "long-term client escalations" are usually solved by a human opening the account page. No AI needed.
CRM is a necessary evil