Everyone just clones a CEO's voice or uses a generic "professional" AI voice. That's lazy and sounds fake. I needed a unique brand voice for security training modules, not a copy of someone else.
I started with a detailed persona sheet: age range, technical background, speech pace, and key emotional tones like "authoritative but not alarmist." I fed this into the Voice Lab, not the cloning tool. Generated multiple samples, rejected anything that sounded like a default news anchor. Fine-tuned by adjusting stability and similarity sliders based on short text prompts that matched our scenario, like explaining a complex firewall rule. The key was using very specific, context-rich prompts that described the *situation* and desired *delivery*, not just adjectives. Took about twenty generations to get it right.
show me the logs
Totally agree on the detailed persona sheet. We used a similar process for a fintech onboarding series. The trick we found was also feeding in a few sample scripts from our actual support team, not just hypothetical prompts. It gave the voice a grounding in real human phrasing patterns.
One caveat, though - did you run the final voice by a fresh audience? We had to dial back the "authority" a notch after testers felt it was a bit too stern, even though our internal team loved it. Sometimes what fits the persona on paper can miss the emotional mark in practice.
Cheers, Henry
"Twenty generations to get it right" is exactly where they get you. Have you calculated the platform cost of those twenty rounds of high-quality inference? That's the hidden labor of this whole "scratch" method. It's still just buying a different flavor of synthetic voice from a vendor's lab.
What's the actual, measurable difference in listener retention between your custom voice and a decent stock one? If you haven't A/B tested that, you've just paid a premium for a feeling.
Your stack is too complicated.