I've been evaluating WellSaid Labs for a potential internal tool, specifically for generating training narrations. As part of the POC, I attempted to create a voice clone of our CEO using a sample script he recorded. The results were... underwhelming. The tone is flat and lacks his characteristic cadence, even though the core timbre is somewhat recognizable.
I followed the basic guidelines: a clean 15-minute recording in a treated space, uploaded via the Voice Cloning portal. The output sounds technically "clear" but completely misses the energy. I'm wondering if I'm missing crucial configuration steps post-upload.
* **Script Selection:** Does the content of the training script matter? I used a dry technical document. Should it be more conversational?
* **Voice Settings:** Are there adjustments to `Stability`, `Style Exaggeration`, or other sliders that significantly impact cloned voice performance? The documentation is vague.
* **Audio Pre-processing:** My pipeline is currently: `ffmpeg` for normalization and noise reduction → upload. Should I be doing more?
Here's the basic `ffmpeg` chain I used:
```bash
ffmpeg -i raw_input.wav -af "afftdn=nf=-20, loudnorm=I=-16:TP=-1.5:LRA=11" output_clean.wav
```
Has anyone successfully cloned a dynamic, expressive voice and gotten beyond a robotic monotone? I'm looking for specific workflow or parameter tweaks, not just "use more data." The audio quality is high, so I suspect the issue is in how the model is tuned.
-- latency
sub-100ms or bust