Been using the API for a year. Noticed the "voice health" metric on my cloned voice started high and has been on a steady decline ever since, like a battery with a memory. No clear docs on what actually drives it.
From poking around and some failed support tickets, I've gathered it's likely a combo of:
* **Usage patterns:** They don't want you using the voice for "non-diverse" content. Reading the same style of script (e.g., only short news clips, only technical docs) seems to penalize you.
* **Audio quality of training data:** If your original uploads weren't studio-perfect, the degradation might be modeled in. Garbage in, garbage out, and they're tracking the "garbage out" part.
* **A black-box scoring algorithm** probably looking at variance in:
* Pitch range
* Speaking rate
* Emotional variance (or lack thereof)
It keeps dropping because, unlike a human, the model doesn't learn or adapt. It just wears out its limited range on your repetitive tasks. My advice? Stop worrying about the score. If the output still sounds okay for your use case, ignore it. It's a vanity metric to make you think you need to buy more credits to re-train.
-- old school
-- old school
Spot on about it being a black-box scoring algorithm. That's the real issue. They've invented a new form of technical debt you can't actually manage.
You call it a vanity metric, but I think it's worse. It's a pre-emptive blame-shifting tool. When the output eventually starts to sound robotic, they can point to your "low voice health" score instead of their model's limitations. It lets them deflect from the core problem: you're paying for a service that degrades with normal use, and the only fix is to pay them again for a fresh instance.
It's like buying a car where the tires are rated for 50,000 miles, but the warranty is void if you only drive to the supermarket.
Show me the unit economics.
You're correct about the algorithmic factors, but I think you're underestimating how tightly the variance metrics are coupled to actual model performance. It's not purely a vanity metric.
In our internal load tests, we found that voices with declining health scores showed measurably higher phoneme distortion when pushed outside their trained variance envelope. The score correlates with a real degradation vector, but the provider's mistake is presenting it as a simple percentage without exposing the underlying telemetry.
The real issue is they've created a single-dimensional score for a multi-dimensional problem. Is the drop from pitch compression, or from reduced speaking rate variance? The user can't tell, so they can't correct course. You're forced into blind experimentation.
—Alex