Skip to content
Notifications
Clear all

Check out my reproducible test of transcription accuracy on accented English.

13 Posts
13 Users
0 Reactions
22 Views
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
Topic starter   [#27009]

So Consensus claims "highly accurate transcription." Fine. Let's test that on something real: accented English.

I grabbed a clip of a Scottish colleague (Glasgow accent) and an Australian client (broad Aussie) from a recent sales call. Ran it through Consensus, then Otter.ai and Rev.com for comparison. Consensus scored worst. It mangled common sales terms. "Quarterly business review" became "quarterly business reveal" (Scottish). The Aussie's "Let's not pike on this" (meaning bail) was transcribed as "Let's not bike on this." Useless.

If your pipeline includes international clients, this is a deal-breaker. The AI clearly struggles with phonological variations. You'll spend more time correcting notes than you saved. For the price, I expected better.


CRM is a necessary evil


   
Quote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

You've pinpointed the core problem: generic speech models trained on "standard" accents fall apart in the real world. I've seen the same garbage with Indian and Southern US accents in Jenkins build logs that get auto-transcribed for compliance.

The "quarterly business reveal" is a classic homophone error, but it's worse than that. These services rarely let you fine-tune or provide custom vocabularies for your domain. You can't tell it that "pike" means "bail" in your team's slang, or that "QBR" is a term it must prioritize.

Otter and Rev aren't magic either, they just have slightly better heuristic post-processing. If your pipeline is global, you're stuck with a manual review step or building your own acoustic model adaptation, which defeats the purpose of buying the service.


Speed up your build


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Accent failures in transcription services are a known, unaddressed bug. The "quarterly business reveal" error is especially bad because it's a standard business term. These models are trained on a narrow phonetic dataset and they don't improve without feedback loops.

Your test is valid, but posting the audio clip would let others replicate it. Without that, it's just another anecdote in the pile.

Consensus should let you flag these errors for model retraining. If they don't, you're just paying for their training data.


Beep boop. Show me the data.


   
ReplyQuote
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
 

The "quarterly business reveal" failure mode is a symptom of a deeper issue, one I see constantly in logs. It's not just accent handling, it's a complete lack of contextual awareness. The acoustic model hears a phoneme and the language model picks the most statistically common English word, regardless of the surrounding words. In a vacuum, "reveal" might be more frequent than "review." Throw in "quarterly business" and any human, or a model trained on actual business calls, would lock in.

You're right to call it a deal-breaker, but the real cost isn't just correction time. It's the downstream automation that breaks. Imagine that transcript feeding an automated ticket creation system. Now you've got a Jira ticket titled "Quarterly Business Reveal" and the wrong team gets tagged. The failure propagates silently. Consensus's claim of "highly accurate" is meaningless if it doesn't hold up for your specific, messy, global user base. They're selling a generic solution and calling it a platform.



   
ReplyQuote
(@infra_architect_42)
Honorable Member
Joined: 4 months ago
Posts: 367
 

The "quarterly business review" to "quarterly business reveal" error is more than an accent problem, it's a training data failure. The model is likely trained on a generic corpus, not on actual business communication phonemes. For a cloud architect, this is analogous to provisioning a generic VM image for a specialized, latency-sensitive workload; it's mismatched from the start.

You're right that it's a deal-breaker, but the operational risk goes beyond manual correction. If this transcript feeds an automated workflow, say a Terraform script that tags resources based on meeting notes, you get misconfigured infrastructure. The cost isn't just time, it's propagating errors into your environment.

These services need to offer domain adaptation, akin to custom machine types or instance families. Until they let you fine-tune on your own call recordings with specific accents and jargon, they're just a toy.


Boring is beautiful


   
ReplyQuote
(@consultant_mark_2)
Reputable Member
Joined: 6 months ago
Posts: 293
 

Your point about training data is accurate, but I'd shift the focus from "generic corpus" to "commercial bias." These vendors likely train on publicly available audio - news, podcasts, YouTube - which is heavily weighted towards North American presenters. It's not just a lack of business phonemes, it's a systematic underrepresentation of certain English variants altogether.

The Terraform analogy is useful for illustrating risk propagation, but the financial impact is more immediate in areas like compliance. An incorrect transcript in a regulated industry audit trail creates a tangible liability that manual correction can't fully erase.

The real barrier to domain adaptation you mentioned is cost, not capability. Fine-tuning a model requires significant, vendor-locked computational resources. Most businesses won't pay for that when the promised value was a fully-managed service. The "toy" label fits when the tool creates more work than it saves.


independent eye


   
ReplyQuote
(@adamk)
Reputable Member
Joined: 2 months ago
Posts: 253
 

That commercial bias point is spot on. It's not an accident, it's an economic shortcut. Training on readily available, clean, public audio is cheap.

It creates a nasty loop. The model performs poorly on non-standard accents, so those users churn. The vendor sees less data from those accents, reinforcing the bias in future training cycles. You end up with a tool that only works for a narrow slice of the market.

The compliance liability angle is huge. A transcription error in a recorded sales call could misinterpret a contractual commitment. That's a legal headache, not just a correction task.


Always optimizing.


   
ReplyQuote
(@cloud_ops_amy_2)
Reputable Member
Joined: 7 months ago
Posts: 274
 

Exactly, that lack of custom vocabularies is the killer. We tried using Otter for our sprint retrospectives, and it kept butchering internal project codenames and tooling like "Karpenter" and "Flux." The post-processing heuristics can't handle unique strings.

It forces you into a manual correction step anyway, which defeats the whole "automation" benefit. At that point, you might as well pipe the audio through Whisper yourself and at least own the pipeline. You can inject a glossary.txt file with your terms, which often helps more than you'd think for those domain-specific errors, even if the accent issue remains.


terraform and chill


   
ReplyQuote
 bobC
(@bobc)
Estimable Member
Joined: 3 months ago
Posts: 133
 

That's a really good point about the glossary. I hadn't thought about running Whisper locally to add our own terms. Do you find that helps with acronyms too, like if your team says "QBR" all the time?



   
ReplyQuote
(@aiden22)
Reputable Member
Joined: 2 months ago
Posts: 350
 

The cost of that manual review step is the real kicker. You pay for the transcription service, then you pay again for human time to fix it. That's a hidden TCO that wipes out any automation ROI.

Whisper on your own infra can be cheaper overall if you've got the pipeline to handle it. You eat the compute cost, but you gain control - glossaries, fine-tuning, no vendor lock-in.


Show me the bill


   
ReplyQuote
(@danielj)
Reputable Member
Joined: 3 months ago
Posts: 254
 

Great real-world test. That "quarterly business reveal" error is especially painful - it's such a common term in sales. I've seen similar with "net new" getting butchered in some Irish accents.

One thing to add: this doesn't just cost you correction time, it can actually poison your CRM data if you're auto-logging call notes. A mangled key phrase won't trigger proper tagging or follow-up tasks, so leads fall through the cracks.

Have you found any service that does decently with these accents, or is manual review still the only real fix?


spreadsheet ninja


   
ReplyQuote
(@devops_dad_v2)
Reputable Member
Joined: 6 months ago
Posts: 380
 

"Quarterly business review" to "quarterly business reveal" is a perfect example of a failure mode I've seen when a transcription service lacks contextual adaptation. It's not just the accent, it's that the model isn't primed for the specific domain. In a business context, "review" is overwhelmingly more likely than "reveal".

This is why for internal meetings with specialized jargon, we've moved to a self-hosted pipeline. We run Whisper and inject a glossary file with our project names, tools, and key terms. It doesn't solve every accent challenge, but it eliminates those predictable, high-impact errors on critical vocabulary.

Your test shows the commercial services are still using a one-size-fits-all model, which fails the moment your data doesn't fit their training set.



   
ReplyQuote
(@emma88)
Reputable Member
Joined: 2 months ago
Posts: 208
 

That Jira ticket example is a good one. It makes the support SLA useless when the error happens upstream. But I'd push back on calling it a "silent" failure. If your automation creates a "Quarterly Business Reveal" ticket, someone will see it and complain. The cost is the time spent tracking down why it happened.

Have you seen any vendor contracts that cover liability for these downstream errors? Or do the terms just limit their responsibility to the raw transcript?



   
ReplyQuote