Hey everyone! I'm super new to voice cloning tools and was really excited to try PlayHT. I've been working on a data analytics tutorial series and wanted to create a consistent voiceover for all the videos.
I followed the instructions for creating a voice clone: I recorded a 5-minute sample of my own voice in a quiet room, reading the script they provided. The sample sounded clear to me! But when I generated the cloned voice, the output... well, it just sounds like a generic, slightly robotic text-to-speech voice. There's zero trace of my actual tone or cadence.
Has anyone else run into this? I'm wondering if I missed a crucial step. My process was basically:
* Uploaded the WAV file (44.1kHz, mono, as per guidelines)
* Let it process for a few hours
* Used the exact same clone for my first test generation
I reached out to support, and they just sent me a link to the general "improving voice quality" FAQ page, which didn't address the clone-specific issue at all 😕. I was hoping for a more detailed walkthrough or maybe a checklist to troubleshoot.
Could you help a newcomer out? I'd love to know:
* What are the most common pitfalls when preparing the initial voice sample?
* Is there a specific format or recording setup (like microphone type) that works best for PlayHT?
* Should I be using different settings when *generating* the audio with the clone, versus creating the clone itself?
I'm really keen to get this working so I can focus on the analytics content itself. Any beginner-friendly advice or recommendations would be amazing! 🙏