Skip to content
Notifications
Clear all

Just built a dashboard comparing Otter's accuracy across different accents.

32 Posts
32 Users
0 Reactions
8 Views
(@ethanc)
Estimable Member
Joined: 2 months ago
Posts: 189
 

Totally feel you on the custom vocab trap - it's such a classic misdirect. The promise of fixing things with a word list feels actionable, but it's solving the wrong problem entirely.

Your point about timing the five-minute clips hits home. We did that exact test and found the correction time wasn't linear. The first minute might take 90 seconds, but by the fifth minute, fatigue and context-switching dragged it closer to 120 seconds per minute. The mental load of deciphering consistent phoneme errors compounds.

That's where the ROI craters. The math assumes a steady correction rate, but the real cost is in the cognitive toll of patching a foundation that's already cracked.


Test, measure, repeat


   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

Your data is solid and matches my own benchmarking. The 90 to low-80s drop for Scottish and Nigerian English is the critical failure point.

You asked about custom vocabulary. It's a dead end for this issue, as it targets nouns, not phonemes. The consistent mangling of consonant clusters you observed is a training data gap in the acoustic model itself. No word list can retrain the phoneme recognition for, say, a Scottish "r" or Nigerian vowel length.

The more interesting test is whether you can use a secondary speech-to-text engine as a fallback for those specific accents. I've had some success piping problematic audio through a locally hosted Whisper model, which sometimes has a different bias, and then merging outputs. It adds complexity but can lift accuracy for those low-scoring accents by 5-7 points.


—Alex


   
ReplyQuote
Page 3 / 3