Hi everyone, I'm pretty new to voice generation tools and have been trying out PlayHT for a project. I'm migrating some old customer service audio guides to a new system, and I'm using PlayHT to regenerate some of the narration.
I've run into a specific issue that's making me a bit nervous about the quality. A lot of the generated speech files have these subtle mouth-click or lip-smack sounds in the pauses between sentences. It's not in every file, but it's frequent enough that I don't think I can just ignore it. It makes the audio sound a bit unnatural and unprofessional.
Could someone guide me through the steps to minimize or eliminate these sounds? I'm using the web interface mostly, and I haven't changed many settings from the defaults. Are there specific voice models that are less prone to this? Or is there a setting for speech smoothness or something similar that I'm missing?
Also, realistically, if I need to process a batch of, say, 50 short clips, what would be the expected timeline to fix this? Would I need to manually edit each one in another program after generation, or can PlayHT handle it on its own? I'm worried this will add a lot of time to my migration schedule.
One step at a time
That's a common issue when you're first getting into voice generation. The good news is you probably don't need to edit each file manually in another program.
Those mouth-click sounds are often artifacts from how the model handles pauses. Before you regenerate everything, try two things in the PlayHT web interface. First, experiment with the "speech rate" or "speed" setting. A slightly faster rate can sometimes reduce the space where those artifacts appear. Second, look for a "pause length" or "prosody" adjustment. Shortening the pauses between sentences can help.
Some voice models are indeed cleaner than others. I'd suggest generating a few test sentences with different "professional" or "studio" tagged voices to compare. If you're still getting clicks after adjusting settings and switching models, then it's time to reach out to their support with a specific example file. They can tell you if it's a known issue with that particular voice.
The suggestions about adjusting speed and prosody settings are a good first step. However, you've asked about the practical timeline for processing 50 clips, which is a crucial project management detail.
If you find a combination of voice model and settings that eliminates the artifact reliably, your timeline is just the batch generation time. If the issue persists even after optimization, you're looking at post-processing. Manually editing 50 clips for clicks, even with spectral repair tools in an audio editor, could add 5-10 minutes per file. That's an additional 4 to 8 hours of work, minimum. Have you factored that potential cost, in both time and possibly software, into your migration schedule?
CostCutter
I ran into the exact same click issue last month. For me, it wasn't just the pause length, but the specific punctuation in my script that triggered it.
Using periods for full stops created longer pauses and more clicks. I started using commas or just letting sentences run together more, and that helped a lot. It made me realize how much the text formatting drives the audio artifacts.
Did you notice if the clicks happen more after certain words or punctuation in your guide scripts?
That's a really sharp observation about punctuation triggering the artifacts. The model's prosody engine definitely parses punctuation as a hard acoustic instruction. A period doesn't just signal a pause, it can trigger a reset in the vocal tract simulation, which is where those glitchy closure sounds originate.
I'd add a caveat though: swapping periods for commas can inadvertently create a run-on, breathless narration style, which might not suit a customer service guide. A more controlled method is to keep the periods for correct grammar, but insert explicit SSML break tags with a specific, shorter duration (like ``). This gives you a clean, consistent pause duration that often bypasses the model's default, artifact-prone pause generation. Not all TTS interfaces expose SSML, but PlayHT's API does, and you can sometimes paste SSML into the web UI.
Have you found that the comma trick works consistently across different voice models, or does it vary?
—Alex
Good point about the SSML. It's the proper fix for this.
One thing I've noticed is that the "break" strength matters a lot. `strength="weak"` or `strength="medium"` often still triggers the click, at least with some of the PlayHT models I've tested. You have to force it to `strength="none"` and then set your own `time` attribute. It's a bit of a hack, but it works.
So instead of just ``, I end up using ``. It overrides the model's default pause behavior entirely.
Run it yourself.
I think you're spot on about the risk of using commas and ending up with a rushed tone. It's a workaround that can undermine the very professionalism you're trying to achieve.
One nuance I've seen is that this artifact doesn't seem to be consistent across all languages or accents within PlayHT, even for the same "studio" tag. A North American English voice might have pronounced clicks where a UK English voice using the same script and punctuation doesn't. It suggests the underlying training data for each voice profile is a bigger factor than the general prosody engine.
Have you found that to be the case? It makes a universal fix tricky.
You've gotten a lot of good technical advice already, especially around using SSML break tags with `strength="none"`. That's likely the most reliable fix if you're able to edit your input text in that way.
But since you mentioned you're mostly using the web interface and are new to this, I'll add a practical, immediate step you can take without diving into SSML just yet. In the PlayHT voice selector, pay close attention to the "Style" tags, not just the "Professional" or "Studio" labels. Look for voices tagged specifically with "Clear," "News," or "Announcer." In my experience, those models are trained on audio that's been meticulously edited to remove breaths and mouth sounds, so the underlying data is cleaner. They often have fewer of these artifacts right out of the gate, even with default punctuation.
It's a quicker test than re-tagging all your scripts. Generate the same problematic sentence with a few different voice styles and compare. You might find one that works for your project tone and sidesteps the issue altogether.
api first