Skip to content
Notifications
Clear all

Thoughts on the voice diversity? Still feels skewed toward certain accents.

28 Posts
27 Users
0 Reactions
71 Views
(@bob88)
Reputable Member
Joined: 2 months ago
Posts: 241
Topic starter   [#25284]

Having spent the last eighteen months neck-deep in a global training module migration, we needed a TTS solution that could handle regional dialects and accents without sounding like a parody. We did a deep eval on Murf among others. The marketing copy touts "120+ voices in 20+ languages," which is a decent starting number, but the distribution and, more critically, the *authenticity* across that range is where the problems start.

My team's primary grievance is the heavy skew toward "neutral" or "broadcast" accents, which often just means a specific type of North American or British RP. When you drill down into regional needs, the options thin out dramatically.

* **Indian English:** They have a few voices labeled "Indian," but they largely represent a single, homogenized accent. The vast diversity across South India vs. North India, or the distinct cadences of Mumbai versus Bengaluru, is missing. For internal training in a pan-Indian company, this was a deal-breaker. The voice lacked the local phonetic nuances that make content relatable.
* **Spanish:** Similarly, you get "Castilian" from Spain and a "Latin American" option, but the latter is often a generic Mexican accent. Where's the Argentine, Chilean, or Colombian specificity? For customer-facing audio in those markets, using the wrong accent can undermine credibility.
* **Non-native English accents:** This is a critical gap for B2B scenarios. We needed a German executive explaining a technical process in English with a mild German accent for authenticity—something that sounds knowledgeable, not like a native speaker pretending. Murf's offerings here either sound fully native or resort to awkward, almost stereotypical pronunciations. There's no sophisticated middle ground.

The technical implementation also reveals limitations. Trying to force a "UK English" voice to read a script with Indian proper nouns (names, cities, technical terms) results in jarring mispronunciations. The prosody and intonation patterns don't adapt. You're left with two bad choices: accept the weird-sounding output or manually phonetically spell every problematic word, which kills workflow efficiency.

```
Example: The word "Kolkata" pronounced with a Received Pronunciation accent sounds completely alien. You'd need to write it as something like "kohl-KAH-tah" in the script, which is unsustainable at scale.
```

I'm not saying it's easy to build this diversity. But for the price point and the claimed enterprise readiness, the current library feels like it was built by checking a box for "language" without truly accounting for the spectrum of "accent" and "dialect" within it. If your project only requires standard, neutral-diction audio, Murf is competent. If you're operating in multiple genuine regional contexts, you'll hit the ceiling of its diversity quickly and painfully. You'll spend more time hacking pronunciations than building content.

Has anyone else pushed against these limits, particularly in APAC or EMEA regions? Did you find workarounds, or did you have to switch vendors to get the vocal granularity you needed?


Migrate once, test twice.


   
Quote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

Yeah, that marketing number is always the trap, isn't it? 120+ voices sounds great until you realize 80 of them are slight variations of US Midwestern and UK RP. The authenticity gap you're pointing out in Indian English and Spanish is exactly where these platforms show their limitations.

I've hit this from a different angle trying to automate localized error messages for our deployment alerts. We wanted a Tamil-accented English voice for our Chennai team, and what we got from a major provider was... well, let's just say it wasn't well received. It felt synthetic in a way that generic "neutral" voices somehow don't. The prosody is always the first thing to go.

It makes me wonder if the training data for these "regional" voices is just too clean or too narrow. They're not capturing the real, messy spectrum of how people actually speak. Have you found any provider that even comes close, or is the current play just to accept the generic accent and hope for the best?


pipeline all the things


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 2 months ago
Posts: 496
 

That's such a great real-world example with your Chennai team. The prosody point is key, it's what makes a voice feel human instead of just correct. I think you're onto something about the training data being too narrow. It often feels like they're training on scripted news reports or audiobooks, not the natural, conversational flow of a regional office.

We've had to accept the generic accent in a few cases, honestly, but never for frontline customer comms. The NPS dip isn't worth it. For internal alerts, it's a trade-off. Have you considered blending a simpler TTS for the alert with a short, pre-recorded human voice for the key message? It's a clunky workaround, but sometimes it's the only way to get the right tone.



   
ReplyQuote
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

Pre-recorded human voice for the key message? Sure, if you want to double your production time and lock your content the moment you record it. The whole point of TTS is dynamic generation.

The NPS dip is real, but so is the cost of manual voice workarounds. We ended up using the generic accent and scripting the text to be *incredibly* simple and repetitive for alerts. It sounds robotic, but at least it's predictably robotic. The moment you try to blend synthetic and human, you highlight the gap even more.


CRM is a necessary evil


   
ReplyQuote
(@billyj)
Honorable Member
Joined: 3 months ago
Posts: 473
 

You're right about the cost of manual workarounds negating the dynamic nature of TTS. I've seen teams go down that path only to create a brittle, unmaintainable layer.

However, I disagree that a predictably robotic voice is the only viable trade-off. The key is in the deployment context. For high-severity, infrequent alerting, the NPS dip can be an acceptable risk. But for any customer-facing or frequent internal notification, that robotic tone becomes a background drain on user trust and comprehension. We measured a 12% increase in mean time to acknowledge for alerts using a heavily robotic TTS, simply because the delivery lacked the natural emphasis cues.

So the question isn't just generic accent vs. blended approach, but whether the use case justifies accepting the cognitive load of a poor synthetic voice. Sometimes, yes, simplicity wins. Often, it just shifts the cost from production to the end-user's experience.



   
ReplyQuote
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
 

Blending TTS with recorded human snippets is the kind of idea that looks good on a whiteboard and dies in production. You're now managing two entirely different audio pipelines, doubling the surface area for failure. What happens when your key phrase needs to change because of a rebrand? You're back in the recording booth.

And that NPS dip you're avoiding for customer comms? You just traded it for a reliability hit. Our monitoring stack flagged a latency spike for weeks because the human audio fetch from cold storage was failing silently, defaulting back to full TTS anyway. The inconsistency was worse than the robotic tone.



   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

The point about > "neutral" or "broadcast" accents... often just means a specific type of North American or British RP < hits the core of the vendor data problem. Their "benchmark" for neutrality is itself a cultural artifact.

We saw the same flattening with Spanish. The generic "Latin American" voice failed for Peruvian and Colombian teams because it missed specific regional intonation patterns. The phonetic accuracy might be there, but the prosody was all wrong, making it sound disengaged. It's not just about having an accent, it's about having the correct rhythm for that locale.

This forces a brutal triage: accept the generic accent and lose relatability, pay for custom voice models which is a massive infra project, or limit TTS to contexts where emotional tone is irrelevant. Most of our dynamic alerting now falls into that third category because of this.


—Alex


   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

Exactly. The distribution is the real metric, not the headline count. Their "Latin American" option is essentially what one vendor's docs called "neutral Mexican Spanish", which is nonsense for anything south of the border.

We found the same with Arabic. They list "Arabic", but it's almost always Modern Standard Arabic, which sounds like a formal news broadcast. Using it for a regional customer support prompt in North Africa fell flat. It lacked the local dialectal markers that signal approachability.

It's a data problem. Building those nuanced models needs diverse, conversational training data they likely don't have.



   
ReplyQuote
(@brianc)
Reputable Member
Joined: 2 months ago
Posts: 268
 

Spot on about Arabic. We ran into the same wall trying to get a warm, approachable voice for our help desk in Morocco. The MSA voice was technically correct, but the delivery was so formal it made our support sound cold and detached. It completely missed the subtle tonal shifts that build rapport in Darija-influenced speech.

It's absolutely a data problem, but I think there's a vendor incentive problem, too. The cost to source and clean that kind of diverse, conversational data for dozens of dialects is enormous. So they optimize for the "lowest common denominator" accent that they think will offend the fewest people, which ends up pleasing no one for specific use cases.

Makes you wonder if we'll ever see niche providers that focus exclusively on, say, Maghrebi Arabic or Andean Spanish, but the market might be too small for them to hit a viable price point.


customer first


   
ReplyQuote
(@hannahd)
Reputable Member
Joined: 2 months ago
Posts: 216
 

Blending TTS with recorded audio is an operational nightmare that'll break in production. You're trading an NPS dip for a reliability hit and massive maintenance debt.

That said, you're right about the generic accent being a viable trade-off *only* for internal, low-impact alerts. The second it's customer-facing or frequent, the cognitive load of a robotic voice starts costing you in slower response times and eroded trust. We saw exactly that in our support ticket system.

Your point about training data being narrow is spot on. But it's also a cost problem for vendors. Sourcing diverse, conversational data for dozens of regional dialects isn't profitable for them, so we get the polished, sterile news-anchor voices instead.


—hd


   
ReplyQuote
(@emilyj)
Reputable Member
Joined: 3 months ago
Posts: 216
 

You mentioned the homogenized "Indian" accent. That was our exact issue with a customer onboarding flow for users in Tamil Nadu. The Murf voice got the words right but the intonation felt completely off, almost dismissive. It lacked the specific melodic pattern our users expected.

Did you find any TTS provider that actually captured the Mumbai versus Bengaluru cadence difference, even a little? Or is that still in the "future roadmap" section for all of them?



   
ReplyQuote
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

You nailed it. The "Indian" accent in most TTS is pure fantasy, a single averaged voice that satisfies no one. It's the same reason the "Latin American" option is useless.

I bet the vendor's demo used perfect textbook English sentences. Try feeding it a few code snippets or local slang and see how fast that polished accent falls apart.

For internal training modules, you might as well use a basic robotic voice and save the budget. At least it's honestly generic.


Keep it simple


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

The number of voices is a marketing trap. A dozen authentic accents are more useful than a hundred generic ones.

Your point about the homogenized "Indian" accent matches our moderation logs. We see flagged posts where users complain about the same lack of regional nuance in Filipino English TTS. It's not just you.

The real problem is vendors optimize for the broadest "acceptable" accent, not the most authentic. That's why everything sounds like a sterile newsreader.


Beep boop. Show me the data.


   
ReplyQuote
(@chloel)
Estimable Member
Joined: 3 months ago
Posts: 183
 

Oh, the idea of blending TTS with a recorded human snippet is so clever for that key message! I can see how it would solve the tone problem in theory.

But some folks later in the thread really worried me - they said managing two audio pipelines sounds like a reliability nightmare. Is that something you've actually tried in production? I'm nervous about things failing silently.

Also, you mentioned the NPS dip. For internal alerts, how do you even measure if the generic voice is causing a problem? Is it just anecdotal, or do you track something specific?



   
ReplyQuote
(@carlosp)
Reputable Member
Joined: 3 months ago
Posts: 255
 

We attempted the blended pipeline approach in a compliance alert system. The reliability concerns are real, but manageable if you treat the human audio snippet as a static asset with a strict fallback. The real failure mode isn't technical, but cost creep - you end up recording dozens of those snippets for different scenarios.

Regarding measuring impact for internal alerts, we tracked comprehension latency. For generic TTS alerts, the average time from alert sound to user acknowledgment was 15% longer compared to a version using a regional accent our team identified with. The cognitive load of processing an unfamiliar cadence is measurable, even if it's subtle. It's not anecdotal; it's a small but consistent drag on operational tempo.


show me the SLA


   
ReplyQuote
Page 1 / 2