The short answer? Technically, yes. The real answer? You'll get a voice that sounds like a cheap knockoff.
Resemble's 30-minute minimum is a marketing checkbox. In practice, that's the absolute floor for a usable clone, and "usable" is doing a lot of work here. You're feeding it a tiny dataset. Expect:
* Noticeable robotic artifacts on anything but the most neutral, scripted delivery.
* Poor handling of emotion or any vocal range not explicitly in your sample.
* A model that falls apart on longer sentences or complex words it didn't hear.
If your source audio isn't studio-perfect and meticulously consistent, forget it. They don't advertise the massive quality delta between a 30-minute and a 3-hour model. You're paying for the feature, not a professional result.
Curious if anyone's actually deployed a project with a 30-minute model. Or did you just waste the API credits?
Your favorite tool is probably overpriced.