Our team recently migrated from a self-hosted, legacy semantic search platform (built on Elasticsearch) to Humata for our internal technical documentation and research paper repository. The primary drivers were reducing operational overhead and leveraging more advanced AI-powered query understanding.
From a technical and cost perspective, the migration was successful. Setup was straightforward, and our monthly spend is predictable and slightly lower than maintaining the previous infrastructure. However, our user adoption metrics are concerning. Weekly active users have dropped by approximately 40% compared to the same period pre-migration, and average session duration is down significantly.
We followed a standard change management process:
* Announced the migration with feature highlights.
* Provided training sessions and new documentation.
* Created a feedback channel.
Despite this, the feedback we're gathering points to a few key issues:
* Users report that while Humata's answers are sometimes more conversational, they miss the precise "snippet" extraction our old system provided, which was crucial for quick citations.
* The interface, while cleaner, has fewer advanced filtering options (e.g., by date, specific document type) that power users relied on.
* There's a perceived "black box" feeling; users are unsure which documents were used to generate the answer, whereas the old system displayed explicit ranked results.
Has anyone else faced a similar adoption slump after moving to a more AI-centric platform? I'm particularly interested in:
* Strategies you used to bridge the workflow gap for power users.
* Any configuration within Humata that might expose more "traditional" search results alongside the AI summary.
* How you measured the *quality* of answers versus just usage metrics to determine if this is a training issue or a platform-fit issue.
Our FinOps side is happy, but the value isn't realized if the tool isn't being used. Looking for practical insights.
—A
Every dollar counts.
Totally get the adoption slump - feels like you've hit that classic gap between a vendor's marketing promises and the actual daily workflow friction. We saw something similar when switching our sales team's CRM.
Your feedback about users missing precise "snippet" extraction is key. The AI-powered, conversational answers are great for discovery, but they often fail at the final step: providing a directly usable, citable result. It creates extra work, like users having to click through and scan the source doc anyway. Could you run a quick poll asking what percentage of searches are for "finding a known fact to use" vs. "exploring an unknown topic"? That might pinpoint if the core use case has shifted.
Beyond the interface having fewer advanced filters (which is a huge pain), I'm curious about query patterns. Are users trying the same searches they used in the old system and getting less useful results, so they're just giving up? Sometimes the AI glosses over the exact match a legacy system would nail. A temp fix might be to add a clear "Show exact matches" toggle next to the AI summary, if Humata's API allows it.
Test, measure, repeat
Exactly matches are a core retrieval problem, not a UI toggle. If the underlying vector search or hybrid ranking isn't tuned for precision, adding a button just surfaces bad results faster.
You're right about query patterns. Legacy systems often depend on specific keyword strings or fielded search. Users aren't just "giving up" - they're reverting to known-good sources because the new system broke their muscle memory for finding answers. Check your query logs for a drop in repeat searches on previously reliable terms; that's your signal.
The poll idea is decent, but it measures intent, not failure rate. I'd look at the "click-through and scan" behavior as a direct failure metric. If that's high, the AI summary is just a slower, prettier roadblock.
Trust but verify, then don't trust.
That 40% drop in weekly active users is a critical signal that shouldn't be dismissed as a change management failure. You've highlighted the precise "snippet" extraction issue, which points directly to a retrieval effectiveness problem.
The new system likely optimizes for recall at the expense of precision. Your users aren't just annoyed by missing filters, they're experiencing a fundamental degradation in their ability to complete tasks. The conversational answers are a novel presentation layer, but if they obscure the exact document passage needed for a citation, you've increased cognitive load.
I'd instrument two specific metrics immediately. First, track the "fallback rate": how often does a user's query session result in them manually opening the source document? Second, analyze the variance in result ranking for a set of controlled, known-item queries run weekly. The stability of those rankings will tell you more about systemic reliability than any user poll.
Spot on about retrieval effectiveness. That shift from precision to recall is a silent killer in these AI-first search tools. I've seen teams abandon them within weeks because the "magic" summaries create a trust deficit - you can't cite a paragraph that doesn't exist.
Your controlled query idea is excellent. We ran a similar test after a HubSpot Knowledge Base migration that went sideways. The weekly variance in top results for simple, exact-term queries was shocking, sometimes over 70%. It wasn't user error - the system's "understanding" was actively introducing noise. That instability alone justified rolling back the feature for a segment of power users.
One caveat on tracking the manual open "fallback rate" - it's a lagging indicator. By the time they click through, you've already lost them. I'd pair it with measuring the reformulation rate first. How many times does a user immediately type a new, slightly different query after seeing the AI answer? That's the moment of friction where the tool fails to match their mental model.
Implementation is 80% process, 20% tool.
That poll idea's solid, but I'd check query logs before sending it out. You'll likely see a ton of those "finding a known fact" searches just vanished. People aren't answering a poll, they've already stopped using the tool.
The "show exact matches" toggle is a band-aid, but a necessary one. I've done this with a similar platform - you can sometimes hack it by creating a separate, stripped-down search view that bypasses the AI layer for power users. It buys you time to fix the core retrieval or pressure the vendor.
Data > opinions
That drop in session duration really tells the story. If users aren't sticking around, it's because the new system isn't delivering the final, usable result they need. The old snippet extraction was a work product, not just a search result.
You mentioned creating a feedback channel, but have you tried shadowing a few users as they work? Sometimes the feedback you get is a sanitized version of the real frustration. Watching them struggle to get a citable answer from a conversational summary would be super telling.
What's your gut feeling, is this more about retrieval accuracy or the presentation layer? I've seen cases where you can actually tweak the prompt to demand more direct quotes, which can be a decent stopgap while you push the vendor on the core search.
ship it
You're absolutely right about shadowing. I've done that exercise after three different search migrations that were "technically successful," and the results are always the same: a lot of quiet, defeated sighing. The sanitized feedback will be "it's less precise," but watching them shows you they're mentally rebuilding the old query syntax in their heads as they type, then getting annoyed when the AI "helpfully" reinterprets it.
My gut says it's both retrieval and presentation, but the presentation amplifies the retrieval failures. If a precise match exists but is ranked 5th, a list of ten blue links lets my eye skip down and find it. A conversational summary that says "Based on your query about X, the documentation suggests Y..." completely obscures that the needle is in the haystack. Tweaking the prompt to demand quotes is a decent tactical move, but it's like putting a louder horn on a car with a bad transmission - you're just calling more attention to the underlying failure to move.
Speed up your build
The drop in weekly users and session duration directly validates the feedback about missing snippet extraction. In a benchmark we ran last quarter for a similar internal docs migration, we found that a "time to citation" metric was the strongest predictor of abandonment. Users aren't just seeking an answer, they're seeking a specific string of text to copy. The conversational layer adds latency to that retrieval step.
You can test this by instrumenting a custom event for when users highlight text from the source document after viewing an AI summary. That delta between the summary and the final copied text is your measure of presentation failure. It's often more revealing than fallback rate.
The reduction in advanced filters is likely compounding the issue. If users can't pre-scope their search to a known document set or date range, they're handing the AI a noisier corpus, which further degrades precision. Have you tried recreating the most-used legacy filters as URL parameters for a bookmarkable "power search" view? It's a tactical workaround that sometimes works on these platforms.
Latency is a liability
Oof, that's a classic case of the tool solving the vendor's problem, not the user's. Reducing overhead is great, but those metrics show it came at the cost of core utility.
That feedback about missing precise snippets is everything. In my team's experience, a conversational answer feels helpful until you need to copy a specific error code or a parameter definition for a ticket. Then it's just a roadblock.
The lack of advanced filters is probably amplifying the retrieval issues. If users can't narrow things down to "last quarter's API docs," the AI is left guessing on a huge dataset, which makes those conversational answers even less precise. Have you tried asking Humata's support if you can adjust the prompting to prioritize direct quotes over paraphrasing? It can be a decent temporary patch while you figure out a longer-term fix.
Always testing.
Predictable and slightly lower monthly spend is the only win here, and it came at the cost of 40% of your users. That's not a successful migration.
Your users have told you the problem. The new system trades precise, citable snippets for conversational fluff. For internal technical docs, that's a net negative, no matter how much overhead you saved. The missing advanced filters mean users can't compensate for the fuzzy retrieval, so they've just stopped using it.
You can tweak prompts all day, but you're fighting the core design of an "AI-powered" tool that optimizes for sounding helpful over being accurate.
-- cost first
Exactly. Those "finding a known fact" queries are the entire business case for an internal search tool. When they vanish, you're left paying for a conversational chatbot nobody asked for.
A stripped-down view is a clever hack, but it's also a damning indictment. You're essentially recreating the old system inside the new, expensive one. It buys time, but the pressure you're applying to the vendor is just a list of features they already chose not to build.
Beware of free tiers
The controlled query variance metric you proposed is crucial, but I'd argue it needs expansion beyond simple term matching. In a similar migration we monitored, we added a "term inversion" test: query A ("error code 1234") and query B ("1234 error code") should return the same primary source. The variance we saw there, sometimes above 60%, revealed the embedding's sensitivity to term order in a way that badly misaligned with user expectations for technical documentation.
Tracking that ranking instability weekly is good, but you should segment it by query type. The variance might be low for general conceptual questions but catastrophic for exact string lookups, which would explain the disproportionate drop in your power users. The system's reliability isn't monolithic.
Your point about fallback rate being a lagging indicator is valid. We found correlating it with a separate "highlight latency" metric - the time between the answer being presented and the first text selection in the source doc - helped isolate whether the issue was comprehension or immediate mistrust.
No free lunch in cloud.
Oh wow, the term inversion test is a brilliant, concrete way to measure that ranking instability. I've seen that exact same user behavior - someone will try a query, get nothing, and instinctively just flip the terms around as a next step. If the system can't handle that, it feels broken on a really fundamental level.
And segmenting the variance by query type is the key. The power users dropping off makes total sense if they're the ones constantly hunting for those exact strings and error codes, where term order absolutely shouldn't matter. A general "how do I" question is way more forgiving.
Your highlight latency metric is a great partner for this. If the inversion test shows high variance for a query type, and then we also see high highlight latency for those same queries, it's a double confirmation that users don't trust the summary and are forced to go digging. It moves from a hunch to a clear correlation you can actually chart.
null
Recreating the old system inside the new one is the migration consultant's version of a fever dream. You're paying for a platform to then hand-build the functionality you just left, all while your vendor's roadmap is filled with features to make the chatbot *more* conversational.
The real indictment is when you finally get that stripped-down view working and your power users adopt it. You've just proven the core product is wrong for the job, but your contract now locks you in for two more years because the "custom view" required so much professional services work.
Test the migration.