Hi everyone! I've been using PlayHT for a few weeks to create voiceovers for my Shopify product explainers. Overall, I really like it, but I've run into something confusing and I'm wondering if I'm doing something wrong.
I'll generate a voice for a video script, let's say using the "Ethan" voice. I'll download it, and it sounds great. Then, the next day, I realize I need to add one more sentence to the script. I go back, use the exact same voice ("Ethan"), the same settings (speed, pitch), and generate just that new sentence. But when I stitch it to the previous audio, the tone is... off. It's clearly the same voice model, but it sounds like it was recorded in a slightly different mood or room. The cadence feels a bit different, and it's really noticeable when played back-to-back.
Is this a common thing? I'm not changing any advanced settingsβjust hitting generate on a new session. Do I need to be saving some kind of seed or session file to keep it 100% consistent? Or is the expectation that you have to generate the entire script in one go for perfect consistency?
It's making my editing process a bit of a headache, as I have to re-generate whole scripts for tiny changes. Any tips or is this just how it works? Thanks for any insight! 😅
Yes, it's a known issue with many TTS engines, not just PlayHT. The "mood" shift you're hearing is likely due to latent space variance between generation sessions. Even with identical text and settings, the model's starting noise seed can differ.
You can't fix it without a seed value. The expectation for broadcast-quality consistency is to generate the entire script in one session and never append later. For edits, you have to re-generate the whole segment from a logical break, not just the new sentence.
Some services offer a "voice stability" or "seed" parameter in advanced settings. If PlayHT doesn't, you're stuck with the full re-gen workflow.
Metrics don't lie.
User634 nailed the technical reason. That latent space variance is the core of it. Their advice to re-generate from a logical break is the standard workaround, but it's a pain.
It exposes a bigger issue with these platforms. They sell you on a "voice," but what they're really providing is a range of outputs from that voice model. True voice consistency is a feature, and many services treat it as a premium one or don't offer it at all. Check if PlayHT has a "voice stability" slider or a "seed" box buried in an advanced menu. If they don't, you're confirming their platform isn't built for iterative, professional editing.
For your Shopify work, you need to plan scripts as final blocks. Any edit means re-doing the whole block. It's inefficient, but that's the current state of most affordable TTS.
βJW