Hey everyone. I'm dipping my toes into AI voice cloning and wanted to share a small project I just finished. I used Resemble AI to clone five distinct customer service agent voices for a demo system.
The basic workflow was: I recorded short voice samples from volunteers (with permission, of course), uploaded them to Resemble, and used their API to generate scripted responses. I plugged the generated audio files into a simple web interface to simulate a multi-agent support hub. The cloning quality was pretty solid for a demo, though I noticed one voice sometimes had a slight metallic edge on longer sentences. Curious if others have tried similar setups for prototyping.
Interesting prototyping approach. The metallic edge on longer sentences might be a data quality issue - not enough sample material for sustained prosody. I've seen similar artifacts when the source audio clips are too short or have inconsistent background noise.
What's your plan for the data pipeline side if you scale this beyond a demo? Managing the voice assets, versioning the audio outputs, and logging API calls becomes a real problem fast. You'll need a solid metadata layer.
garbage in, garbage out
Resemble's pricing gets steep fast for production. Five voices on their Pro tier plus API calls? That demo cost is about to multiply. What's your expected monthly audio generation volume?
always ask for a multi-year discount