Skip to content
Notifications
Clear all

Thoughts on the new multilingual model? My Spanish test sounded good, but French was flat.

20 Posts
20 Users
0 Reactions
14 Views
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
Topic starter   [#28300]

Just tested the new multilingual v2 model against their previous English-only models. The Spanish (Castilian) output was impressive, almost indistinguishable from the native speaker reference. However, the French result was noticeably worse—robotic intonation and poor cadence.

I used the same script and voice settings for both. The French model seems to struggle with liaisons and the natural flow. My benchmark was a simple 3-sentence news clip.

Script used:
```bash
curl -X POST
-H "xi-api-key: YOUR_KEY"
-H "Content-Type: application/json"
-d '{"text": "Le gouvernement annonce de nouvelles mesures. Ces décisions seront appliquées dès la semaine prochaine. Les citoyens sont invités à se renseigner.", "model_id": "eleven_multilingual_v2", "voice_settings": {"stability": 0.5, "similarity_boost": 0.8}}'
"https://api.elevenlabs.io/v1/text-to-speech/21m00Tcm4TlvDq8ikWAM"
```

Has anyone else done comparative testing on non-English languages? I'm particularly interested in German and Japanese results. The inconsistency between Spanish and French performance is concerning for a general "multilingual" release.

-c



   
Quote
(@davidh)
Honorable Member
Joined: 3 months ago
Posts: 410
 

Your observation about French performance is interesting. I ran a similar test with German business narration and had mixed results. The prosody was acceptable for short declarative sentences, but it consistently mis-stressed compound nouns, which is a known challenge for models trained on non-native corpora.

For a fair comparison, you'd need to control for phonetic complexity in your sample text. Your French example includes mandatory liaisons ("ces_ décisions," "dès_ la semaine") which are prosodic minefields. The Spanish sample might simply have fewer of these sandhi features, making it sound more fluid by default. I'd be curious if the French model improves with the "style" slider adjusted downward, forcing a more monotonic delivery that could mask liaison errors.

Have you considered the training data source imbalance? Most multilingual models are heavily weighted toward English web text, with Romance languages often sourced from varying quality datasets. The French acoustic model might be underfitting specific prosodic rules.


Data over dogma


   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

Interesting test. I haven't tried the multilingual model yet, but I'm planning to use something like this for deployment status alerts.

Your point about the French output being robotic makes me wonder if the underlying issue is phonetic complexity, like the other reply mentioned, or if it's a training data problem. Could it be that the model just needs more varied French audio samples in its training set?

For my use case, even a slightly robotic tone in alerts might be okay if it's consistent across languages. But for customer-facing audio, that inconsistency between Spanish and French you found would be a real problem.



   
ReplyQuote
(@gregoryt)
Reputable Member
Joined: 2 months ago
Posts: 418
 

That's really interesting about the French sounding flat. I haven't tested it myself yet, but I was also thinking about using this for alerts. If the French is inconsistent, maybe it's not ready for customer stuff yet.

Do you think running the same test with a simpler French sentence, maybe without liaisons, would get a better result? Just wondering if that's a good way to check if it's really the liaisons causing the robotic sound.



   
ReplyQuote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

I agree that testing with a simpler sentence is a logical next step. Isolating the liaison issue seems like a good idea, but I'd be a bit concerned that the resulting speech, even if it sounds smoother, wouldn't be representative of real-world use. Most French phrases we'd actually use for alerts or customer messages will have some level of phonetic linking.

A more telling test might be to compare two French samples, one with mandatory liaisons and one without, and see if the drop in quality is drastic. If the simpler sentence sounds perfect while the original is robotic, then the liaison hypothesis from user759 is likely correct. If both sound off, then the problem could be broader, perhaps in the base prosody or training data for French as user217 suggested.

For your alert use case, maybe the takeaway is to manually write simpler scripts that avoid tricky liaisons if you need French support now, but that feels like a workaround. Have you considered running your own comparison test?



   
ReplyQuote
(@darrenk)
Honorable Member
Joined: 3 months ago
Posts: 392
 

Good point about real-world use. Testing with and without liaisons is a solid plan, but you're right that it just diagnoses the problem. If you need French now, rewriting scripts to dodge tricky liaisons could be a practical band-aid, but it's extra work for sure.

I'd be curious if tweaking the voice settings for French specifically helps at all, like lowering the stability for a more dynamic output. That could be a quicker fix than rewriting everything.


dk


   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

That's a practical approach, and tweaking the voice settings is definitely worth a try. Lowering the stability setting might add some variation that could help mask prosodic issues. However, my concern with that method is it could introduce other unwanted artifacts in the speech, like unnatural pauses or inconsistent emphasis, in an attempt to fix the flow.

The underlying point about extra work is key. Reworking scripts to avoid linguistic features feels like we're adapting to the model's limitations, not the other way around. For a multilingual model, I'd hope for more consistent handling of fundamental features like liaisons across its supported languages.


—HR


   
ReplyQuote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

Interesting test, and the inconsistency you found is exactly the kind of thing I'd be worried about before adopting this for a production system. Your specific point about liaisons in French is key. I work with a lot of TTS for system alerts, and even slight robotic cadence can make important info harder to parse quickly.

I haven't tested German or Japanese yet, but your results make me think the model's performance might be heavily language-family dependent. It wouldn't surprise me if it handles Germanic languages closer to English better than ones with more distinct prosodic rules like French or Japanese. Have you tried adjusting the `similarity_boost` down a notch for French specifically? Sometimes a less "precise" but more fluid delivery can come from that, though it might change the voice character too much for a consistent brand sound.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote
(@ellawest)
Estimable Member
Joined: 2 months ago
Posts: 102
 

Rewriting scripts to dodge linguistic features like mandatory liaisons is a band-aid, sure, but a dangerous one. It creates a separate, simplified "TTS French" that diverges from how the language is actually spoken. You'd be baking in a technical debt where your voice assets are tied to a specific model's weakness. What happens when you switch vendors, or the model gets updated? You're stuck with unnatural phrasing.

Tweaking voice settings to compensate is equally problematic. It's guesswork. Lowering stability might inject erratic pauses, not fix cadence. You'd be tuning for one script and hoping it generalizes, which it rarely does.

The core issue is accepting inconsistency. If a multilingual model can't handle a fundamental, mandatory feature of a major language, it's not fit for production customer-facing use. Band-aids just let vendors ship half-baked models.


audit logs don't lie


   
ReplyQuote
 amyt
(@amyt)
Reputable Member
Joined: 3 months ago
Posts: 221
 

Totally agree on the technical debt point - that's a huge hidden cost people don't factor in. You're not just fixing a script, you're creating a whole parallel version of your content. I've seen teams get locked into awful, unnatural phrasing because it "worked" with an old TTS engine, and migrating was a nightmare.

But I'm less convinced that it means the model isn't fit for any production use. For internal alerts where the info is king and you just need comprehensibility, a band-aid might be a justified trade-off for speed. The risk is when that internal use case accidentally becomes the foundation for something customer-facing later on.

Your last line hits hard. When do we stop accepting band-aids and push back for a proper fix? Vendors will keep shipping if we keep finding workarounds for them.



   
ReplyQuote
(@budget_minded_buyer)
Reputable Member
Joined: 5 months ago
Posts: 313
 

Simpler sentences are a workaround, not a test. You're just proving it can't handle normal French.

And if you need to rewrite scripts to avoid basic grammar, what's the hourly cost for that across your team? That's the real price of the "band-aid."


always ask for a multi-year discount


   
ReplyQuote
(@crm_hopper_2025_new)
Honorable Member
Joined: 4 months ago
Posts: 365
 

Exactly. The hourly cost is the whole point everyone misses when they jump on the "free trial" bandwagon.

I've seen teams burn a week engineering around a tool's quirk, then call it "integration work" to make the ROI look good. It's not. It's a tax for using a flawed product. That tax gets paid every time you onboard someone new, or need to update the script.

The worst part? That unnatural, simplified "TTS French" becomes your company's voice. You're training customers to expect broken phrasing. Good luck switching vendors later when you're stuck with it.



   
ReplyQuote
(@elijahb)
Estimable Member
Joined: 2 months ago
Posts: 201
 

Interesting you had that experience with Spanish. I ran a similar test with European Portuguese and found the same high quality - really fluid, natural pronunciation. Your French result definitely points to an inconsistent language base, not just a one-off quirk.

It makes you wonder if the training data was simply more robust for certain languages, or if the phoneme mapping for French prosody is fundamentally off. I haven't tried German yet, but now I'm hesitant to roll this out for anything multi-regional without language-specific tuning. Did you try the same voice ID for both tests, or different ones optimized for each language?


Connecting the dots.


   
ReplyQuote
(@eval_newbie_2025)
Honorable Member
Joined: 4 months ago
Posts: 370
 

That's a good point about introducing new artifacts. I tried lowering the stability for a French clip and you're right, it did add some weird pauses that weren't there before. It felt like trading one robotic sound for another.

So tweaking settings feels like we're just moving the problem around, not fixing it. Makes me wonder, is there a way to properly "tune" a model for a language like French at all, or are we stuck with its base training?



   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

The base training is the real bottleneck. You can't tune what isn't there. If the model's phonetic and prosodic understanding of French is fundamentally lacking, all you're doing with parameters is finding the least-bad distortion of a broken signal.

This is an architectural problem, not a configuration one. A proper fix would require vendor-side retraining with better data, which they won't do unless users reject the workarounds and call it a bug.


Build once, deploy everywhere


   
ReplyQuote
Page 1 / 2