That's an insightful point about training data. The generic tech docs versus messy negotiation conversations is likely the core issue. You can train a model to recognize "SLA" as a term, but without the conversational context of how it's actually defined or contested in a call, it becomes a token to be swapped with similar tokens.
Your suggestion to benchmark against a pre-transcribed recording is the most efficient path. It moves the analysis from "is it accurate?" to "where is it inaccurate, and at what cost?" If you map the errors, you'll likely find they form a pattern around conditional phrases and numerical modifiers, not just single terms. It's the difference between "95% uptime SLA" and "95% uptime SLO" that the model completely misses, because the statistical difference in its training corpus is negligible, while the contractual difference is material.
p-value < 0.05 or bust
Your skepticism is well-founded, and the thread's responses confirm your preliminary tests. While I haven't deployed Sembly specifically for audit trails, I have a parallel observation from testing it against Datadog's incident postmortems.
> The real-world WER in complex, jargon-heavy vendor discussions
The WER on pure jargon is often deceptively low. The systemic failure is in the connective tissue around it. In a discussion about error budget allocation, Sembly perfectly captured "burndown rate" but inverted the conditional logic, transcribing "if the budget is consumed" as "once the budget is consumed." That changes the contractual posture entirely. For audits, this is worse than a misheard word; it's a misinterpreted intent, and it's not reflected in a standard WER score.
The JSON export does contain a session ID and a processing timestamp, which is adequate for chain-of-custody metadata. However, the utility collapses when the timestamps are only provided at the paragraph level. If you need to verify who said "the penalty applies after 72 hours," and that clause is buried in a three-minute transcript segment with four speakers, the immutable metadata is a moot point. You're forced back to the raw audio, which defeats the purpose of a searchable transcript.
Given your requirements, a formal benchmark will likely just quantify the failure modes you've already identified. A more telling test would be to measure the "critical clause error rate" separately from the overall WER, focusing solely on phrases containing dates, monetary values, and scope modifiers.
Your preliminary test results are the benchmark you're looking for. You've already identified the failure modes that matter for audit integrity: diarization collapse with >4 participants and hallucination of critical terms.
The JSON metadata is there, but it's a technical checkmark. The real issue is that the transcript's accuracy is non-uniform. Errors cluster on the very data points that create audit liability, like dates and monetary values. A tool with a 5% overall WER but a 25% error rate on numbers is functionally useless for your purpose.
Running a formal benchmark will just quantify the magnitude of a failure you've already qualitatively proven. If you must proceed, design your test corpus entirely around conditional statements and numerical modifiers from real contracts, not generic meeting chatter. That's the only WER that matters for compliance.
FinOps first, hype last
You've put your finger on the operational reality: quantifying a failure you've already observed is a form of procurement theater. The 5% overall WER versus 25% error on numbers is a perfect illustration of a tool being optimized for the wrong metric, like a database touting high read throughput while its write consistency is abysmal.
If you're forced to build that benchmark corpus from real contracts, go one step further. Don't just test "five days." Test ambiguous phrasings that are common in calls, like "within five to seven business days" or "not to exceed twenty five thousand." That's where these models often splice clauses from different speakers or drop the modifiers entirely, creating a factual statement from a conditional one. The error clustering isn't random; it's structurally biased against precision.
You're right to be skeptical of the marketing, and your preliminary tests are telling you everything you need to know. The issues you've seen in other tools - diarization collapse with more than four people, hallucinations on key figures - aren't edge cases for Sembly; they're the core weaknesses of this entire category when you apply it to audit-grade work.
The immutable metadata is there, but that's the easy part. It's like getting a signed certificate of authenticity for a document where the text inside changes every time you read it. The real-world WER will look great on a spreadsheet, but the errors will be catastrophically clustered on the exact phrases you need to be immutable. I've seen it swap "shall not" for "shall" in indemnity clauses, turning a prohibition into a requirement.
If you're still required to run a benchmark, focus it on those specific failure modes. Don't test general accuracy; test the accuracy of numbers, dates, and conditional language from your actual past contracts. That'll give you the data you need to push back or, if you must proceed, build in a mandatory human review layer for any clause containing a figure or a deadline.
Raise the signal, lower the noise.
Yep, that speaker diarization issue with just four people is the deal-breaker. For audit, you need to know *who* said *what*, and if that's broken, nothing else matters.
The metadata is there in the JSON, but like others said, it's just a wrapper. The real problem is the error distribution. It'll nail the generic flow but mangle a key modifier like "up to" or "not less than." You can't audit on that.
Your preliminary test is the benchmark. If it failed there, it'll fail in production for the same reasons.
data over opinions
I haven't used Sembly for audit trails, but your specific question about WER on jargon-heavy calls got me thinking about a quick test you could run yourself. Instead of a full benchmark, maybe just feed it a pre-recorded, already-transcribed 10-minute snippet from a past vendor call that has dense jargon and conditional language. Align the outputs manually.
What you'll likely see is that the WER on standalone terms is fine, but the model stumbles on the logical connectors between them. Something like this:
```python
# Example of what the model might miss
actual = "The credit is issued provided the SLA is not breached."
transcribed = "The credit is issued if the SLA is breached."
```
That single-word swap changes the entire clause, and it won't show up as a high WER. The metadata will be pristine, but the content is legally flipped.
Clean code, happy life
Thanks for sharing this concrete example. That swap from "provided... not" to "if" is exactly the kind of silent failure that's so dangerous for audits. The intent gets inverted in a way that's easy to miss on a quick review.
Your suggestion to test with a short, pre-transcribed snippet is practical. It bypasses the need for a huge benchmark and isolates the exact problem. Maybe even focusing on sentences with "unless" or "subject to" would show similar flipping.
Do you think using a short, known-good audio clip like that could become a standard smoke test for any transcription tool meant for contracts?
I haven't used Sembly for this, but your concerns about speaker diarization and hallucinated numbers are spot on. I ran a similar tool against some vendor calls last year, and the transcript looked great until you hit the money parts. It turned "escalates to 12% after 90 days" into "escalates to 20% after 19 days." The JSON had all the right timestamps, but the content was fiction. That's not an audit trail, it's liability.
My two cents? The smoke test idea from user542 is the way to go. Take a two-minute clip with a few conditional phrases you've already transcribed, run it through Sembly, and see if it flips the logic. If it does on your known sample, you've got your answer without a full benchmark.
it worked on my machine
That "procurement theater" line is painfully accurate. You see teams burn cycles building elaborate benchmarks just to validate what a simple smoke test already revealed.
Your point about ambiguous phrasings is exactly where these tools falter. I'd add that the structural bias is often toward simplification. The model hears a complex, legally-weighted phrase and outputs a cleaner, more generic version, stripping out the precision that creates auditability. "Not to exceed twenty five thousand" becomes "up to twenty five thousand," which feels similar but carries a different contractual weight. The error report won't flag it as critical, but it changes the obligation.
—daniel
I agree about the core weakness being the error distribution, not the overall WER. The immutable metadata is a technical checkbox, but it's meaningless if the content is wrong.
You're right to focus on jargon-heavy calls, but in my last test the bigger issue was cross-talk and overlapping speech in those negotiations. Even with three participants, the moment two people verbally agreed over each other with a "yeah-yeah," the diarization broke and attributed the following key clause to the wrong speaker.
A simple smoke test with a pre-recorded clip of overlapping agreement followed by a conditional statement will likely show you more than a full benchmark on clear audio.
Totally agree about your preliminary tests being the best indicator. If other tools failed on speaker diarization and critical language, Sembly will likely have the same issues. The tech stack underneath is often similar.
You asked about the immutable metadata in exports. It's usually there, but as others have said, it's just a wrapper around flawed content. The real-world WER on jargon-heavy calls is misleading because the errors aren't spread out. They concentrate on the key conditional phrases you actually need to audit. A tool can have a 95% accuracy and still completely invert a liability clause.
Your smoke test idea is the right move. Take a two-minute clip of a past negotiation with overlapping speech and a tricky "subject to" clause, run it through their free trial, and see what breaks. That'll give you your answer faster than any benchmark.
Docs save time
Your preliminary tests mirror my own migration headache last quarter. The speaker diarization fell apart with just three people once side conversations started. More than a nuisance - it's a complete audit fail.
On the metadata question, the exports do include meeting IDs and timestamps, but that's just a fancy label on a broken product. The real killer for jargon-heavy calls isn't the WER percentage, it's *where* the errors land. I had it transform "the penalty is waived if sub-clause 4(b) is satisfied" into "the penalty is waived after sub-clause 4(b) is satisfied" on a finance call. One preposition change, a totally different financial obligation.
Skip the full benchmark. Take your most complex two-minute clip with overlapping "yeah, go aheads" and a few "notwithstanding" clauses, and run it through their trial. That'll give you your answer in 30 minutes. If it flips the logic there, you can't trust it for the full negotiation.
Mobile vs desk phone is a side channel. The real diarization killer is pricing plan.
Check if the speaker confusion happens more on their cheaper per-user tier. Some vendors silently throttle diarization accuracy on lower plans. You get four speaker labels, but they're essentially random after the first two minutes.
"BAU" to "B.A.U." is a classic tell. It means the model was trained on generic business audio, not procurement or compliance calls. You're paying for a generalist tool and hoping it works for a specialist need. That's always a bad bet.
always ask for a multi-year discount
The mobile vs desk phone comparison is interesting, but my experience points to the network stack as a bigger factor than the handset. Poor VoIP compression on a conference line can strip the high-frequency cues the diarization model uses for separation, making both sides sound similar regardless of device quality.
Your note on predictable errors is key. When "Type 2" consistently becomes "type too," it reveals a training data gap in compliance terminology. That's a fundamental mismatch for an audit trail tool. A high error rate on random filler words is noise; consistent errors on key terms make the transcript unusable.
I've seen diarization degrade with any background noise or echo, which is often worse on mobile in a non-silent environment. But the core confusion between vendor and internal roles usually stems from the model having insufficient speaker enrollment time at the start of a call, not the audio codec.
CloudCostHawk