I’ve been using ElevenLabs for a specific, consistent workflow since their earlier days, generating short-form narration with a particular voice profile. When v2 rolled out with all the fanfare about improved expressiveness and stability, I was expecting a net positive.
My experience has been the opposite. The "improved" model seems to have regressed in a key area for me: consistency of tone within a single generation. With v1, I could feed it a clean, neutral script and get back a predictable, steady delivery. V2 introduces these weird, almost imperceptible shifts in cadence and timbre mid-sentence, especially on longer sentences. It feels like it's trying to be more "natural" but ends up sounding slightly unsettled, like the speaker is distracted.
I ran a simple paired comparison using the same scripts and voice clone from before and after the switch. The difference isn't subtle when you're listening for it. The v2 output has a higher rate of unnatural stress on prepositions or odd pauses that weren't in the script. It's as if the pursuit of prosody introduced a new form of artifact.
Anyone else doing systematic comparisons, or are we all just accepting the vendor's benchmarks? I'm curious if this is isolated to certain voice profiles or if it's a broader trade-off they made. For my use case—where consistency is more valuable than dramatic inflection—this feels like a step backward.
Data skeptic, not a data cynic.
Oh man, that "pursuit of prosody" line hits home. I've seen the exact same pattern in other AI services, where a "quality" bump actually breaks a specific, predictable output you've built a workflow around.
Your point about unnatural stress on prepositions is spot on, that's the kind of subtle artifact that ruins a professional flow. It reminds me of when a text-to-speech engine I used for accessibility docs started randomly injecting dramatic pauses... in the middle of compound nouns. The new "expressive" model was technically "better" on their metrics, but completely unusable for my case.
Have you tried tweaking the stability slider way up? Sometimes that can reign in the unwanted "expressiveness" and get you back closer to the v1 monotone, though it's a band-aid.
it worked on my machine
That's a really solid parallel case, the accessibility doc one. It perfectly illustrates the core tension in these platform updates: what's statistically "better" on broad benchmarks can be catastrophic for a niche, repeatable workflow.
Your band-aid suggestion with the stability slider is a good first step, and it's often what support will recommend. From my experience managing these transitions on the enterprise side, though, it's rarely a full fix. You're often just trading one artifact for another, like introducing a flat, robotic dullness instead of the weird cadence shifts. The underlying prosody model has fundamentally changed.
It puts users in a tough spot, having to hack around a "quality" improvement. Have you found any other services where pulling back an "expressiveness" control actually gave you the old model's behavior, or is it always a different flavor of compromise?
Architect first, buy later
Your paired comparison is the right methodology, and you've hit on the key issue: vendor benchmarks almost never measure consistency of tone within a single generation. They're looking at aggregate MOS scores or similarity across many clips, which can completely miss the micro-artifacts you're describing.
We saw this with a TTS model update last year. The "improved" version scored higher on expressiveness in listener surveys, but our internal metrics tracking spectral centroid and pitch variance over time showed new instability. The model was overfitting to dramatic training data, causing exactly those distracting mid-sentence shifts.
Have you quantified the "higher rate of unnatural stress"? A simple count of misplaced lexical stresses per 100 words in your v1 vs. v2 outputs would turn your subjective observation into a concrete regression metric you could report.
Show me the benchmarks
The real issue here is that "improved" on the roadmap usually just means "optimized for new customer acquisition." They're chasing the metrics that look good in sales demos and on review sites, not the ones that matter for a mature, repetitive workflow. That shift in cadence you're hearing is the model trying to perform for a first-time listener, not deliver consistent utility for the hundredth job.
Your paired comparison is the right move, but good luck getting vendor support to care. Their success metrics are probably all about onboarding satisfaction and churn in the first 90 days. Someone who's built a stable, predictable pipeline around v1 is already captured revenue. The noise from users like you gets drowned out by the applause from new sign-ups who think the expressiveness is "cool."
It's the classic enterprise bait-and-switch, just dressed up as an AI model update. You're not paying for the service you had, you're beta-testing the service they want to sell to the next wave.
Your mileage will vary