Skip to content
Notifications
Clear all

ELI5: How does the 'voice health' score work, and why does mine keep dropping?

3 Posts
3 Users
0 Reactions
8 Views
(@devops_rookie_2025)
Reputable Member
Joined: 2 months ago
Posts: 203
Topic starter   [#7399]

Hi everyone! I'm still pretty new to WellSaid Labs and trying to figure out all the features. I've been using a custom voice for a few weeks, and I keep seeing its 'voice health' score slowly go down.

Could someone explain how that score works in simple terms? Like, what actually makes it decrease? I want to make sure I'm using the tool correctly to keep my voice sounding good. Thanks in advance for any help! 😊



   
Quote
(@lindae)
Estimable Member
Joined: 6 days ago
Posts: 54
 

Ah, the classic "health score" mystery. It's a staple of SaaS platforms to give you a vague, anxiety-inducing metric that inevitably trends down, encouraging you to either buy more credits or upgrade your plan.

In simple terms, it's likely a composite score based on how often you use the voice, the variety of your inputs, and maybe some internal model metrics on output consistency. It decreases because the system is designed to flag voices that aren't being actively "trained" with diverse data, or perhaps to nudge you toward generating more content to keep their usage numbers up. The real question isn't how to keep the score up, but whether a slight drop actually correlates with any perceptible loss in audio quality for your specific use case. Have you noticed your voice actually sounding worse, or are you just worried about the number going down?


Trust but verify.


   
ReplyQuote
(@cloud_cost_breaker)
Estimable Member
Joined: 2 months ago
Posts: 131
 

That's a good question. It's not a SaaS dark pattern like the other reply suggests. Think of it like a statistical model's confidence interval.

Your custom voice is a machine learning model trained on your initial samples. The "health" score is essentially a measure of prediction confidence. Every time you generate audio, the system compares the output against the model's expectations for that voice. If you're only feeding it very similar, short, or repetitive text inputs, the model has less varied data to reinforce its patterns, so its confidence in producing consistent, high-quality output for a wider range of inputs decreases. The score drops.

To maintain it, you need to periodically generate more varied content - different sentence lengths, emotions, and phonetic combinations - to give the model reinforcing data. It's less about "training" and more about preventing model drift.


Less spend, more headroom.


   
ReplyQuote