Skip to content
Notifications
Clear all

Why is Otter.ai so inaccurate with non-native accents?

7 Posts
7 Users
0 Reactions
2 Views
(@crmsurfer_43)
Estimable Member
Joined: 5 months ago
Posts: 102
Topic starter   [#20793]

I've been testing Otter.ai for our revops team to automate meeting note-taking, especially for our discovery calls with global prospects. While it's fantastic for my North American colleagues, I've noticed a significant drop in accuracy when our team in Singapore or our clients in Eastern Europe are speaking.

Just last week, a call with a prospect from Poland had the transcript littered with nonsense words. Key product names and pain points were completely garbled. It turned a potentially useful record into something we had to manually re-listen to, defeating the purpose. I've seen similar issues with strong accents from India, Nigeria, and Scotland within our own company.

I'm curious if others have run into this. From a technical standpoint, I assume it's a training data issue—the models are likely optimized for General American or RP English. But for a tool marketed to global teams, this is a major workflow blocker. We can't have a "tiered" accuracy system based on where someone is from.

Has anyone found effective workarounds or settings adjustments? Or is this just a known limitation we have to accept? I'm comparing notes with a colleague who uses it for HubSpot sales call logging, and they're facing the same hurdles. It makes me wonder about the underlying speech recognition engines and whether competing services handle this diversity any better.



   
Quote
(@averyt)
Eminent Member
Joined: 5 days ago
Posts: 21
 

Oh, absolutely. You've nailed the core issue with your training data point. It's a known pain point across most speech-to-text services, honestly. The models are overwhelmingly trained on specific, "clean" accent datasets.

A workaround I've had moderate success with is to use Otter in tandem with a more robust recorder, like a dedicated hardware mic, to get the cleanest possible audio feed into the tool. Even that only helps so much, though. For truly critical calls, we sometimes run a second, parallel transcription using a different service that's marketed as more multilingual, like Sonix, and then merge the best parts. It's a clunky extra step, but it beats a completely garbled transcript.

Have you tried adjusting the speaker labeling settings at all? Sometimes manually tagging a speaker with a region hint *before* the meeting can nudge the engine slightly.


Automate all the things


   
ReplyQuote
(@amandaf)
Estimable Member
Joined: 1 week ago
Posts: 73
 

You're right about the training data, and your workaround is a common but frustrating reality. Running two services defeats the whole value proposition of automation and simplicity.

The "nudging the engine" tip with speaker labels is, frankly, a band-aid on a much larger problem. It shouldn't be on the user to pre-tag speaker demographics for basic accuracy; that's the vendor's job to handle from the raw audio. If Otter and others want to serve global business teams, their training datasets need to reflect that reality. Marketing a product as a meeting assistant while it fails on a standard international call is a mismatch.


—AF


   
ReplyQuote
(@emilya)
Estimable Member
Joined: 1 week ago
Posts: 75
 

Agreed, but the mic suggestion is minimal gain. The signal-to-noise ratio improvement is maybe 5-10% at best for these models. The real bottleneck is the acoustic model itself.

Your parallel service workaround just proves the point. If you're paying for a tool, you shouldn't need a redundant second tool as a crutch. It's a vendor problem, not an audio quality problem.

Manually tagging a speaker region shouldn't be necessary. The model should infer that from the audio stream. Expecting users to pre-label accents is asking them to do feature engineering for a SaaS product.


Prove it with a benchmark.


   
ReplyQuote
(@auditor_abby)
Estimable Member
Joined: 4 months ago
Posts: 111
 

It's a vendor problem, but the contractual one is often overlooked. That signal-to-noise ratio gain is irrelevant if the SLA for accuracy is based on a narrow dataset they won't disclose.

You pay for a tool expecting it to work for your documented use case. If your vendor contracts and SOC 2 reports don't specify the linguistic and acoustic diversity of their training data, you're accepting a black box with no recourse. The real cost isn't the second tool, it's the labor for manual correction and the compliance risk of inaccurate records.

Ask them for the demographic breakdown of their model training sets. Their inability to provide it is your answer.


Where is your SOC 2?


   
ReplyQuote
(@fionah)
Estimable Member
Joined: 1 week ago
Posts: 80
 

Finally someone brings up the paper trail. The SLA loophole is the whole game.

You can ask for the demographic breakdown, but good luck. They'll cite proprietary models and trade secrets. The real question is whether their marketing claims of being a "meeting assistant for global teams" constitute a material misrepresentation if their training data is demonstrably narrow.

I've seen procurement teams push for a contractual accuracy minimum per defined accent group, with a testing protocol. It usually gets laughed out of the negotiation. That tells you everything about their confidence in serving a non-U.S. market.


trust but verify


   
ReplyQuote
(@devops_dad_joke)
Estimable Member
Joined: 4 months ago
Posts: 104
 

You've hit the nail on the head about the procurement laugh test. It's the perfect signal.

We tried that exact playbook a few years back with a different "global" transcription vendor. Their sales engineer got real quiet, then said their SLA was for "general accuracy" and that "demographic-specific guarantees were not industry standard." Which, yeah, that's the whole problem, isn't it?

The irony is, my team in Manila ended up building a shadow system with Whisper models we fine-tuned ourselves on local call recordings. It's not as polished, but it works better for our actual team. Kind of funny when the paid SaaS product gets sidelined by a cobbled-together open source project.



   
ReplyQuote