Skip to content
Notifications
Clear all

Hot take: The 'stability' and 'similarity' settings are a black box. Need more docs.

4 Posts
4 Users
0 Reactions
31 Views
(@data_pipeline_ops)
Reputable Member
Joined: 6 months ago
Posts: 176
Topic starter   [#21643]

I've been experimenting with the API for a small project to generate synthetic voiceovers for data pipeline explainers. I keep hitting a wall with the 'stability' and 'similarity_boost' parameters.

The official docs are pretty thin. Changing the values often gives me unexpected results—like a voice becoming more robotic when I increased stability, which I thought would do the opposite. I'm basically just guessing at this point.

Has anyone done any systematic testing or found a good rule of thumb for these settings? A concrete example of a good combo for a clear, neutral narration would be a huge help.

Building my first pipeline.


PipelinePadawan


   
Quote
(@clarag)
Reputable Member
Joined: 3 months ago
Posts: 274
 

Totally feel your pain on this. I was also surprised when higher stability made a voice sound more flat and lifeless. I've been sticking with `stability=0.4` and `similarity_boost=0.75` for my project's tutorial narration. It seems to hit a sweet spot of consistent but not robotic, at least for my ear.

Maybe it's less about finding a universal 'good' setting and more about adjusting for the specific voice? Some of the pre-made voices already have more expression built in. What voice are you using for your explainers?



   
ReplyQuote
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 467
 

> Maybe it's less about finding a universal 'good' setting and more about adjusting for the specific voice?

That's likely part of it, but without real docs we're all just burning credits to find out. I'd bet the underlying cost model for inference changes with these parameters, too. Higher stability might be using a cheaper, more deterministic path.

If you're sticking with 0.4 and 0.75, you should at least do an A/B test against the default settings. Run the same script, see if your character/second cost changes. Otherwise you might be paying a premium for a setting you settled on by accident.


show me the bill


   
ReplyQuote
(@devops_journeyman)
Reputable Member
Joined: 5 months ago
Posts: 216
 

Yeah, that tracks with my experience. I had the same misconception about stability - I thought higher meant smoother delivery, but it seems to lock the voice into a narrower, more predictable pattern.

For clear narration, I've had good results with fairly low stability (around 0.3) and a moderate similarity boost (0.6). This seems to keep the core voice recognizable while allowing a bit more natural inflection. Try pairing that with the `Sarah` voice if you're on ElevenLabs; it's been pretty reliable for technical stuff.

Have you noticed any difference in generation time when you tweak these?



   
ReplyQuote