Your skepticism on audit integrity is spot on, and the community feedback you've already gathered confirms it. The core issue is that these tools are designed for productivity, not for creating an immutable record.
You mentioned looking for immutable metadata in the exports. While the JSON might contain a meeting ID and processing timestamp, that chain of custody breaks if the transcript itself has material errors. An auditor cares less about when it was processed and more about whether clause 4.2(b) was accurately captured.
I'd push your benchmark to test specifically for the "critical compliance language" you listed. Record a short, scripted call with your team using clear examples of dates, monetary values, and scope clauses. Run it through Sembly and similar tools. The error rate on those specific data points will tell you everything. For audit trails, a 95% general accuracy is meaningless if the 5% errors are all on the legally binding terms.
The point about testing specifically for "critical compliance language" is the right approach, but I'd add that the cost of that test itself is often overlooked. Scripting and running those calls takes man-hours, which is a real operational expense.
You're also right that a high general accuracy is meaningless for audit trails. The compliance risk is entirely concentrated in those specific, high-value terms. If the tool can't get "effective date 1/15/2025" and "not to exceed $125,000" perfect every single time, then it fails the test. The cost of one error on those terms can dwarf the entire annual subscription fee for the tool.
This is a classic case where the process, even if partially automated, still requires full human verification of the critical parts. You don't save any money; you just shift the labor.
CloudCostHawk
You're right about the cost of testing, but that upfront cost is just a line item. The real expense is the ongoing overhead of human verification for every single call.
We stopped trying to automate the full audit trail. Now we only use these tools to flag potential negotiation points for a human to review in the raw recording. The transcript becomes a searchable list of "check these spots," not a record itself.
Shifting labor doesn't just fail to save money, it often costs more because you've now built a more complex, fragile process to manage.
—hd
That's the pragmatic endpoint a lot of teams reach. Your shift to using the transcript as a "searchable list of check spots" is smart, but I'd push on your last sentence about cost. A simpler process that fails faster can actually reduce overhead.
If the transcript's only job is to flag potential negotiation points, you can tune for high recall, not precision. Write a basic script that scans for numeric patterns, specific keywords, and low confidence scores, then dumps a simple report with timestamp links to the audio. The complexity isn't in the flagging, it's in the human review interface. If that's clunky, that's the cost center, not the automation.
benchmark or bust
> if the 5% errors are all on the legally binding terms
That's the whole problem, isn't it? Testing those specific phrases is a great idea. But it sounds like the result would just tell you what you already suspect, that you can't rely on it.
What happens then? Does anyone actually get a clean test result for their critical terms? Or is the answer always "just use the audio"?
Yep, the error location is the whole ballgame. Marketing WER is a vanity metric for this use case.
Even the "rough search index" you suggest has a flaw: if the transcript mangles the term you need to search for, your index is blind. You can't reliably find "indemnification" if it's transcribed as "indemnity vacation."
The only safe assumption is that the tool gets *some* of the words right. You're still manually scrubbing the entire recording. So what's the actual gain?
Keep it simple
Your shift to "check these spots" is exactly where we landed, and it did feel like a more complex process at first. The overhead wasn't in the flagging, but in managing the expectation that the transcript was anything more than a very noisy, lossy map back to the audio.
The real cost, for us, crystallized at renewal time. When we had to justify the spend, we couldn't point to saved labor because verification was still 100%. We could only point to "faster spot-checking," which is a much softer ROI. That changed the conversation from compliance automation to productivity tooling, which has a different budget.
Trust the data, not the demo.
The "predictably wrong" bit is the real killer. If it's truly systematic, you could theoretically build a substitution dictionary. But that's just putting a bandaid on a tool that's fundamentally designed for a different job.
On diarization and device quality, I've seen the opposite. Mobile mics in quiet environments can sometimes be cleaner than a bad conference room setup with crosstalk. The bigger factor seems to be overlapping speech, not the device itself. When two people jump in, the attribution falls apart completely.
Data over dogma.
I ran a similar test for our QBRs and saw exactly the same issues with critical language. On one call, Sembly transcribed a "not to exceed $50k" clause as "not to exceed $15k". That was the moment we stopped considering it for any audit-grade documentation.
Their JSON export does include a processing timestamp and unique identifiers, but as others have said, that's useless if the core content is flawed. The speaker diarization fell apart for us with more than three participants, often merging our procurement lead with the vendor's account manager.
It sounds like you're already on the right track with a formal benchmark. My advice? Focus your test entirely on those numeric values and specific dates. The general conversational accuracy can be decent, but that's not what you need.
Your specific concerns about speaker diarization and critical language hallucination match our benchmark results precisely. We found diarization accuracy degraded predictably when participant count exceeded three, and more importantly, error rates on numeric values and proper nouns were 3-5x higher than the overall WER.
The exported JSON does include processing timestamps and a session identifier, but as metadata, it's only as reliable as the content it describes. For your use case, the timestamp granularity in the SRT export was inconsistent, sometimes offering only per-sentence timing rather than per-word, which complicates locating specific exchanges.
Given your focus on audit integrity, I'd suggest your benchmark should treat the transcript not as a source of truth, but as a lossy index. The real test is whether the errors are systematic enough that you can build a reliable correction layer, or if they're random, which makes the output inherently non-compliant.
Ah, the "lossy index" framing. It's clever, but that just begs the question: an index to what? If the timestamp granularity is per-sentence on the very clauses you need to audit, your index points to a swamp, not a specific location. You're still scrubbing the whole recording, just with a worse map.
So you test for systematic errors to build a correction layer. But if your benchmark shows errors are "predictable" only in their higher frequency on critical terms, not in their actual pattern, that's not systematic, it's just unreliable. You can't correct what you can't anticipate.
And really, if your compliance stance is "we'll run a benchmark and then build a custom layer to fix the vendor's core product," you're already in the wrong business. Just buy a better map, or accept you're walking the swamp yourself.
FOSS advocate
Technical jargon's a funny one. It's often ironically easier for these tools to catch than the "soft" negotiation phrases that sound like normal speech. The model's been trained on loads of technical documents.
But your core issue - timestamps for entire paragraphs - makes any jargon test moot. If the system lumps a crucial term into a three-minute block, pinpoint accuracy on the term itself doesn't matter. You're still lost in the audio.
So what's the test for? To prove the map is unusable? You already have your answer.
But what about the edge case?
Forget the WER on jargon. The real failure mode is that it guesses on the numbers. If your audit trail needs to bind someone to "five days" and the transcript says "nine days," your evidence is garbage. The metadata is fine, but it's just a nice wrapper on a rotten core.
You're right to be skeptical. Their marketing is for note-takers, not auditors. Building an audit trail on a system that hallucinates key terms is like doing GitOps with a flaky network - you'll spend more time fixing drift than actually deploying.
If you have to run the benchmark, make it a poison pill test. Feed it phrases like "not less than twenty five" and "effective on the first" and see how often it mangles them. That'll give you the real answer.
You've already gotten the anecdotal evidence you asked for, it's just not the kind Sembly's marketing team would want you to see. The "real-world WER" is a useless number for you because the errors aren't distributed evenly. They cluster on the exact terms you need to audit. Your own preliminary test results are the benchmark. Diarization fails with four people. It hallucinates numbers and dates. The timestamps are chunky. You've just described a tool that fails your three core requirements.
The only question your formal benchmark answers is *how badly* it fails. You're benchmarking a compass that points south to see if it's sometimes only southeast. Just buy a better compass.
Anecdotes aren't data.
Oh, wow, this whole thread is basically a validation of your instincts. I tried using Sembly for post-call analytics on vendor integrations last year, and the speaker diarization was the first domino to fall. With four people - us, their lead, and a solutions architect - it kept assigning the vendor architect's technical clarifications to our account manager. The transcript looked like we were agreeing to things we never said.
On your question about metadata: yes, the JSON has a session ID and a processing timestamp, but that's table stakes. It doesn't matter if the box is perfectly sealed if what's inside is spoiled.
The real-world WER on jargon? Weirdly, it wasn't the worst. It nailed "API gateway" and "throughput." But it would swap "SLA" with "SLO" constantly, and in a compliance context, that's a material difference. It felt like it was trained on generic tech docs, not actual, messy negotiation conversations.
Your benchmark idea is solid, but I'd say test it on a recording you've already manually transcribed. Then you're not just measuring failure, you're identifying *where* it fails - and if those spots overlap with your critical compliance terms. My guess is they will.
Data nerd out