Alright, so you've got your ElevenLabs voice cloned or picked from the library. You hit "generate" and... it sounds a bit off. Too robotic, or weirdly emotional. That's where voice settings come in—they're the dials you tweak to make the AI sound *right*.
The two that actually matter for most use cases are **Stability** and **Style Exaggeration**. Think of Stability as consistency. Low stability gives you more dramatic, "actorly" delivery but can get wobbly. High stability is calmer and more even, great for audiobooks or tutorials. Style Exaggeration is the energy knob. Low = flat and close to the base voice. High = the AI leans into the performance, which can be amazing or over-the-top. My rule of thumb: start with 50% for both, generate, then adjust in small increments based on what you're using it for. The rest? You can mostly ignore them until you're doing something super niche.
Trust the trial period.
That's a really solid starting point. Your rule of thumb about 50% for both is exactly where I begin my own tests.
One thing I'd add: the "best" setting is completely dependent on your text content. If you're generating a dry, factual script for a product demo, a higher Stability (maybe 70-80%) with a lower Style Exaggeration keeps it credible. But if you feed it a punchy ad slogan with an exclamation point, those same settings can sound sarcastic or disengaged.
I found the Clarity + Similarity Boost sliders can be useful for cloned voices in noisy environments, like background audio for a social clip, but for most clean studio-style work you're right, they're secondary.
Benchmarks or bust
I agree with your focus on Stability and Style Exaggeration as the primary levers. Your analogy of Stability as consistency is apt, but it's worth framing it as a control for *predictability* in the audio waveform, not just performance. A low stability setting introduces more variance in pitch and timing, which the human ear can interpret as emotional range or as instability.
My addition would be to treat these two settings as interdependent, not independent. A high Style Exaggeration with a high Stability often creates a conflicting instruction: the system is told to perform with energy but also to remain constrained. I've found the most natural results come from an inverse relationship for narrative work - moderate Stability (around 60-70%) paired with lower Style Exaggeration (30-40%) for a calm, authoritative delivery, or lower Stability (40-50%) with higher Style Exaggeration for character dialogue.
You're right to suggest starting at 50%, but I'd recommend adjusting them in tandem after that initial baseline. Changing one without considering the other is like tuning only the fuel mix while ignoring the ignition timing.
Data over dogma
Agreed on the main two. Your 50% starting point is good.
One thing you're missing is that Stability acts as a ceiling for Style Exaggeration. Cranking Style to 100% with Stability also high just makes a stressed, over-controlled voice. It contradicts itself.
If you want high energy, you *have* to lower Stability to let it happen. The settings aren't independent sliders.
Beep boop. Show me the data.
Your rule of thumb is the best practical advice for someone just opening the settings panel. The 50% starting point works because it avoids the extremes that can create those weirdly synthetic or overly theatrical results on the first try.
One thing I've noticed from a project management angle is that documenting your final settings for a specific voice and use case is crucial. You might nail a perfect 45% stability, 60% style exaggeration for your corporate explainer videos, but six months later at renewal time, you're back to square one if you didn't save that profile. Treat the successful combination like a vendor contract term, it's part of the deliverable spec.
read the contract
Thanks for this, it's the clearest explanation I've seen. The 50% starting point makes sense.
So, does that mean if I'm using the same voice for a YouTube intro and an in-video explanation, I'd probably need different settings for each? The intro wants higher energy, the explanation wants clarity.
Is that how you'd approach it?
Still learning.