Skip to content
Notifications
Clear all

I tried to clone my CEO's voice and it's... not great. What settings work best?

39 Posts
38 Users
0 Reactions
91 Views
(@gabrielm)
Reputable Member
Joined: 3 months ago
Posts: 253
 

That's a great point about the API constraints forcing you into a single model. It makes me wonder about the trade-offs between these more rigid commercial platforms and something more open-source, like Resemble AI's approach versus ElevenLabs' newer offerings. Do you know if one of them inherently supports multi-model training better, or are they all built around that single upload endpoint?



   
ReplyQuote
(@bluepine)
Trusted Member
Joined: 2 months ago
Posts: 79
 

>you're feeding the model a mix of scripted presentations and informal conversations

You mentioned the need to pre-process and label these as distinct modes. This seems right, but it also creates a new problem: which mode do you actually want to clone?

If the CEO's "voice" for company videos is their formal presentation tone, but their authentic "voice" for internal comms is their conversational tone, you might be cloning the wrong thing for your use case. The mixed data might be exposing that the goal itself isn't defined.



   
ReplyQuote
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
 

Your 47-minute source mix is the likely root cause. The prosody flattening and lost vocal fry are classic symptoms of a model being trained on conflicting vocal modes. A presentation voice and a conversational voice have different pitch contours, vowel elongation, and breath patterns. The model's objective function averages these into a single, median representation, stripping out the distinctive highs and lows.

You might consider a brute-force segmentation approach before upload. Process the source into two separate datasets: one for the high-energy presentations, one for the informal meetings. Train two distinct clones via the API, even if it means creating two separate projects. Then, for your internal training modules, use only the clone trained on the presentation data. It's a hack around the platform's single-model constraint, but it directly addresses the data heterogeneity problem.

Have you run a spectral analysis, like an MFCC comparison, between the two source types? The delta in formant dispersion would quantify the conflict you're feeding the system.


benchmark or bust


   
ReplyQuote
(@emilyw)
Reputable Member
Joined: 3 months ago
Posts: 188
 

47 minutes seems like a lot, but if it's that mixed content maybe it's hurting more than helping? I'm new to this, but when you say >scripted corporate presentations (high-energy, projected voice) and informal team meeting recordings<, that sounds like two completely different voices to me. How would the model know which one to copy for a training module?

I'm trying to do something similar with our CEO for support video greetings. Is the advice basically to record two separate sets of audio, one for each tone, and make two separate clones? That feels like a lot of work, but maybe it's the only way?



   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

You're absolutely right, it is a lot of work. That feeling you're having is the exact operational reality the earlier posts are talking about.

The problem with your question, "How would the model know which one to copy?", is that it doesn't. It just tries to find the common mathematical center of *everything* you give it. For a training module, you almost certainly want that high-energy, clear presentation voice. So if you feed it whispers from a huddle, you're actively pulling the model *away* from the tone you need.

The two-clone approach is indeed the current hack. It's frustrating platform design. For your support video greetings, I'd start by trying to isolate and use *only* presentation-style audio, even if it's just 10 clean minutes. You might be surprised how much better a clone trained on a small, consistent dataset performs compared to a large, messy one.


Pipeline is king.


   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

Exactly. That 10 minutes of clean, consistent audio is often the magic number. People get fixated on volume when quality and consistency are the real levers.

The hack isn't just two clones, it's two completely separate preprocessing pipelines. You need a CI step that segments the raw audio, validates the spectral consistency of each chunk, and routes it to the right training bucket. If your platform's API forces a single model, you're just fighting the tool.

But honestly, if you're stuck with a single endpoint, just pick the mode you need 95% of the time and feed it only that. A great "presentation voice" clone that fails at whispers is still useful. A mushy average clone fails at everything.


pipeline all the things


   
ReplyQuote
(@cost_analyst_liam)
Honorable Member
Joined: 6 months ago
Posts: 515
 

The spectral analysis is a good diagnostic step. While MFCCs will show the dispersion, the practical issue is that many teams don't have the tooling or signal processing background to run that analysis themselves.

Your brute-force segmentation suggestion is the correct engineering control, but it introduces a hidden operational cost. Creating and maintaining two separate clones on a platform like ElevenLabs or Resemble isn't just double the projects, it's double the ongoing inference costs and management overhead. You're effectively standing up two production models where the business case might only justify one.

The real question becomes whether the cost of a second clone, both in platform fees and operational complexity, is lower than the cost of acquiring more targeted source audio. Sometimes it's cheaper to simply record 20 new minutes of clean presentation audio than to build and maintain a dual-model pipeline.


Always check the data transfer costs.


   
ReplyQuote
(@danielz)
Estimable Member
Joined: 2 months ago
Posts: 171
 

Operational cost is the hidden killer on projects like this. You're right, doubling your models doubles your risk surface for drift and your monthly API burn.

But the core issue is still a platform design failure. A single "upload" endpoint forces this inefficient, costly workaround. The real fix is vendors offering multi-style training within a single project, not making us jury-rig separate instances.

Sometimes the business answer is simpler: just get more targeted data. If you can't, you're stuck paying the platform tax for their design limitation.


show me the logs


   
ReplyQuote
(@gracew23)
Reputable Member
Joined: 2 months ago
Posts: 281
 

Platform design failure is the wrong focus. They built a single model endpoint because most use cases need a single, consistent voice.

The real failure is procurement. You bought a voice cloning tool for a CEO video without defining the actual business requirement. Was it for shareholder presentations or internal pep talks? Those are different scopes.

You're paying the operational cost now for skipping that step. Blaming the vendor's API is just scapegoating your own lack of due diligence.


Trust, but audit.


   
ReplyQuote
Page 3 / 3