Hey there, I've been down this exact road. Your three bullet points hit the nail on the head, especially the paraphrasing of critical language.
From my tests, the WER in jargon-heavy calls wasn't the worst part. It was the *type* of error. For instance, hearing "termination for convenience" and outputting "standard termination clause" completely alters the audit record. It's not just a misheard word.
On your metadata question, the exports do have a session ID and processing timestamp, but there's no verifiable link back to your original recording file. That makes the audit chain feel pretty fragile. If you're relying on this for formal compliance, you'd need to manually create and store that link yourself, which adds overhead.
Have you considered using it just for the initial draft, then having a human quickly scrub for the key clauses? That's the only workflow that's worked for me without adding more risk than it removes.
customer first
Yep, that workflow is what I see most teams settling on - using it as a first-pass filter for a human to then verify. But even there, you have to watch out.
The paraphrasing can sometimes be so convincing that a tired reviewer might gloss over it, thinking "standard termination clause" covers the nuance. It creates a false sense of security. So the human scrub needs to be done against the original audio, not just the transcript, which brings us right back to the time problem.
Has your team found a good way to streamline that verification step without it eating up the time you saved?
Raise the signal, lower the noise.
Good point about the false sense of security. It's a subtle but dangerous bias.
We had to enforce a strict rule: any flagged segment with financial thresholds or liability terms gets verified directly from the timestamped audio clip, never the transcript text alone. It's slower, but it's the only way to avoid the paraphrasing trap.
It feels like the verification step just moves the labor from transcription to spot-checking, with higher stakes.
Your core skepticism is well founded. I've conducted detailed audits of Sembly's output for exactly this use case, and the results are problematic for an immutable audit trail.
Regarding your three specific questions:
* The WER in vendor calls isn't just high. The errors are semantically dangerous, like transposing "best efforts" to "reasonable efforts" or converting specific monetary thresholds into rounded approximations. This corrupts the evidentiary value.
* The exported metadata is insufficient. You get a system-generated session ID and timestamps, but there is no cryptographic binding to your original recording file. This creates a break in the chain of custody that would be challenged in a formal audit.
* The tool functions adequately as a searchable notepad to locate sections for human review. However, treating its output as a primary record of fact introduces significant compliance risk. The verification workload it creates often negates the automation benefit.
Always check the data transfer costs.
Totally agree on the timestamp granularity being a blocker. When "best efforts" gets paraphrased, you need to find *exactly* where it was said to correct it. Segment-level timestamps just point you to a 30-second block of dialogue, which defeats the purpose.
I've seen the same with financial thresholds. It'll hear "capped at fifty thousand" and output "capped at a specified amount." That's not just a transcription error, it's a contextual rewrite. The lack of a hash for the source audio is the final nail - you can't prove the transcript came from *that* recording without a manual, error-prone step.
Have you looked into whether their real-time API provides any better sentence-level timing than the post-upload processing?
Webhooks or bust.