Skip to content
Notifications
Clear all

Beginner question: What does 'stability' and 'similarity' actually do in the voice clone settings?

1 Posts
1 Users
0 Reactions
0 Views
(@data_analyst_2025)
Reputable Member
Joined: 3 months ago
Posts: 167
Topic starter   [#23406]

Hi everyone! 👋 I'm pretty new to the whole AI voice space and have been experimenting with PlayHT's voice cloning for a small data presentation project. It's been fun, but I've hit a wall with two specific settings in the advanced options.

When I'm fine-tuning a cloned voice, I see sliders for "Stability" and "Similarity." I've played with them a bit and I can hear the output change, but I don't really understand *what* the engine is actually doing when I adjust them. The tooltips are a bit vague for a beginner like me.

Could someone walk me through, in practical terms, what these two settings control? For example:
* Does **Stability** affect how consistent the emotion or pitch is across the entire generated speech?
* And **Similarity** – is that strictly about matching the original speaker's timbre, or does it include speaking style too?
* If I wanted a voice that sounds very close to my sample but with a more consistent, calm delivery (less dramatic), how would I adjust these?

A real-world example from your own workflow would be incredibly helpful. I'm coming from a data viz/analytics background, so I'm used to tweaking parameters in tools like Looker, but this audio domain is new to me!

Thanks in advance for helping a newbie out. Looking forward to learning from your experiences.



   
Quote