I've been editing a podcast and Descript's Overdub feature looks amazing. I need to fix a few mumbled words without re-recording.
But the tech behind it makes me a bit nervous. Can someone explain in simple terms how it actually clones a voice from a sample? More importantly, where does that voice data go? Is it used for anything else, or stored in a way others could access?
I'm trying to understand the privacy and security side before I train a voice model, especially since I'm working with guest voices sometimes.
Great question, especially when dealing with guest voices. The basic idea, as I understand it, is they use a neural network trained on your voice sample to learn how your specific vocal patterns fit together. It's not just stitching recorded syllables, it's generating new speech that matches your timbre and rhythm.
On the data side, you should check their current privacy policy. I remember reading a while back that when you train a Standard Overdub voice, the model stays on your device. But for a more accurate Professional voice, the training might happen on their servers. That's a key distinction for privacy. If the data is processed on their servers, you'd want to know how long they keep the raw audio you uploaded for training, and if it's used to improve their general models.
Have you looked into whether you can get the needed accuracy with a Standard voice trained locally? That might ease the guest voice concern. I'd be nervous uploading someone else's voice without explicit permission too.