Hi everyone, I'm pretty new to voice generation tools and have been trying out PlayHT for a project. I'm migrating some old customer service audio guides to a new system, and I'm using PlayHT to regenerate some of the narration.
I've run into a specific issue that's making me a bit nervous about the quality. A lot of the generated speech files have these subtle mouth-click or lip-smack sounds in the pauses between sentences. It's not in every file, but it's frequent enough that I don't think I can just ignore it. It makes the audio sound a bit unnatural and unprofessional.
Could someone guide me through the steps to minimize or eliminate these sounds? I'm using the web interface mostly, and I haven't changed many settings from the defaults. Are there specific voice models that are less prone to this? Or is there a setting for speech smoothness or something similar that I'm missing?
Also, realistically, if I need to process a batch of, say, 50 short clips, what would be the expected timeline to fix this? Would I need to manually edit each one in another program after generation, or can PlayHT handle it on its own? I'm worried this will add a lot of time to my migration schedule.
One step at a time
That's a common issue when you're first getting into voice generation. The good news is you probably don't need to edit each file manually in another program.
Those mouth-click sounds are often artifacts from how the model handles pauses. Before you regenerate everything, try two things in the PlayHT web interface. First, experiment with the "speech rate" or "speed" setting. A slightly faster rate can sometimes reduce the space where those artifacts appear. Second, look for a "pause length" or "prosody" adjustment. Shortening the pauses between sentences can help.
Some voice models are indeed cleaner than others. I'd suggest generating a few test sentences with different "professional" or "studio" tagged voices to compare. If you're still getting clicks after adjusting settings and switching models, then it's time to reach out to their support with a specific example file. They can tell you if it's a known issue with that particular voice.