Alright, I’ll bite. I’ve been running Resemble AI through its paces for the last quarter, mostly for synthetic voiceovers in customer onboarding sequences. The sales pitch, as we all know, is all about "human-like emotional range" and "granular control." Forgive my sardonic chuckle.
My specific gripe—and I’m genuinely curious if I’m the only one who feels like they’re taking crazy pills—is with the so-called 'excited' inflection. I’ve fed it pristine source audio of someone who sounds genuinely thrilled about, say, winning the lottery. I’ve toggled the emotion intensity slider to its maximum setting. I’ve even tried the granular prosody edits, nudging pitch and speech rate into what should be the stratosphere of enthusiasm.
The result? A voice that sounds like a mildly caffeinated accountant reading a quarterly earnings report. There’s a slight uptick in pace, a tiny wobble in pitch, but the core *affect* is utterly flat. It lacks that breathless, dynamic, almost unpredictable quality of real excitement. It’s as if their model interprets "excited" as "slightly faster and marginally higher," not as a fundamental shift in vocal weight and timbre.
I’ve benchmarked this against a few other vendors (won’t name them here, don’t want this flagged as promotion), and while none are perfect, Resemble’s emotional palette feels peculiarly muted. It’s fantastic for neutral, authoritative, or even somber tones. But ask it to convey joy, urgency, or sarcasm? It lands with the emotional impact of a soggy paper towel.
So, I’m throwing this to the community:
* Is this a training data issue? Are they just using too many bland corporate voice samples?
* Am I missing some hidden config or a pre-processing trick? I’ve played with the raw audio quality of my source, the script phrasing itself—everything short of sacrificing a goat to the audio gods.
* Or is this simply the current ceiling for their tech? The marketing copy sure suggests otherwise, but we all know how that goes.
I’d love to hear your workflow specifics if you’ve managed to crack this. What are you putting in your source? Any particular phonemes or script structures that trigger a better response? Or are we all just waiting for the next model update, hoping our "excited" voices won't sound like they're discussing a mild interest in spreadsheet formatting?
—Bella
Price ≠ value.