Skip to content
Notifications
Clear all

I tried to clone my CEO's voice and it's... not great. What settings work best?

39 Posts
38 Users
0 Reactions
90 Views
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
Topic starter   [#25525]

I have been conducting an extensive evaluation of WellSaid Labs' voice cloning capabilities for a specific professional use-case: creating a high-fidelity digital voice avatar for our CEO to be used in internal training modules. The goal was to achieve a level of naturalness and character accuracy that would be indistinguishable from a real recording in a controlled, quiet environment.

After several cloning attempts and rigorous A/B testing against source material, the results are, frankly, suboptimal. The output consistently exhibits a noticeable flattening of prosody, a loss of subtle vocal fry present in the original, and an artificial "ring" in sustained vowels. This is not a critique of the platform's overall capability, but rather an analysis of my specific methodology. I suspect the issue lies in my configuration and input data preparation.

Here is a detailed breakdown of my process and the parameters I've manipulated:

**Source Data Preparation:**
* **Duration:** 47 minutes of clean, studio-quality audio.
* **Content:** A mix of scripted corporate presentations (high-energy, projected voice) and informal team meeting recordings (conversational, lower dynamic range).
* **Processing:** I applied the following preprocessing chain using FFmpeg before upload:
```bash
# Normalize to -24 LUFS, apply a gentle high-pass filter at 80Hz, and denoise.
ffmpeg -i input.wav -af "loudnorm=I=-24:TP=-1.5:LRA=11, highpass=f=80, afftdn=nf=-25" output.wav
```
* **Segmentation:** I provided the data as both a single file and as pre-segmented 3-minute chunks, observing no significant difference in the final model.

**WellSaid Labs Configuration Attempts:**
1. **Attempt 1 (Default):** Used the platform's default settings. Result was overly smooth and generic.
2. **Attempt 2 (Enhanced Expressiveness):** Enabled "Enhanced Expressiveness" and "Dynamic Range Boost" in the avatar settings. This introduced unnatural pitch variation, making it sound forced.
3. **Attempt 3 (Manual Tone Adjustment):** Set the base tone to "Authoritative" and speech rate to 90%. This improved the perceived intent but further compromised timbral accuracy.

**Benchmarking Method:**
I performed a Mean Opinion Score (MOS) test with a panel of 7 employees who regularly interact with the CEO. They rated naturalness and similarity on a scale of 1-5.
* **Source Recording (Control):** MOS 4.7
* **My Best Clone Attempt:** MOS 3.1
* **A Generic Corporate Male Avatar:** MOS 2.8

The marginal gain over a generic avatar is concerning given the resource investment.

My core question for the community is multi-faceted: For those who have achieved high-fidelity clones (MOS > 4.0), what was your precise workflow?

* Is there a definitive sweet spot for **source audio duration and composition**? Is pure, consistent speech better than varied emotional samples?
* What is the impact of **audio preprocessing**? Should one avoid any filtering and upload raw, clean audio?
* Beyond the surface-level sliders, are there **advanced techniques in the Lab** or via the API (e.g., custom dictionary emphasis, phonetic tuning) that yield disproportionate improvements?
* How significant is the **text script** provided to the voice during synthesis? Does using sentence structures and vocabulary highly specific to the speaker improve the output?

I will be running another batch of experiments next week, and I am eager to incorporate community findings into my testing matrix. My hypothesis is that the current AI model prioritizes clarity and generalizability over replicating unique, imperfect vocal signatures, but I hope to be proven wrong with the correct configuration.



   
Quote
(@emilyl)
Honorable Member
Joined: 3 months ago
Posts: 527
 

Oh wow, that's a really detailed process! I've been curious about voice cloning for some project update videos, but this level of analysis is way beyond me.

You mentioned using a mix of scripted presentations and informal meetings. Could that be part of the problem? I'd think the different energy levels might confuse the model a bit. Would it be better to stick to just one style for the source audio?

Also, 47 minutes sounds like a lot. Is there a chance too much data makes it harder for the software to pick up the unique character traits? I'm just guessing here, but I'm fascinated by this stuff.



   
ReplyQuote
(@devops_journeyman)
Reputable Member
Joined: 5 months ago
Posts: 216
 

Interesting point about the mix of presentation and meeting audio causing confusion. I think you're right, but it might be less about "energy" and more about microphone distance and processing.

For my own experiments, I got the cleanest results by using a single, consistent recording source. A 20-minute scripted read of neutral material (like a company handbook section) recorded on the same headset mic worked better than an hour of varied sources. The compression and EQ applied to a formal presentation recording can create a different vocal profile than a raw meeting track.

Have you tried isolating *just* the headset audio from those informal meetings? It might give the model a more consistent baseline, even if the sample length is shorter.



   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

You're on the right track with consistency. Different recording setups aren't just about energy, they introduce subtle compression, EQ, and noise profiles that the model has to try and separate from the actual voice. It's a data quality problem.

On the duration, more isn't always better. A clean, consistent 15-minute sample can outperform an hour of noisy, varied material. Think of it like training any ML model, garbage in, garbage out. The software needs to identify the speaker's core timbre, not learn multiple versions of it.

What recording setup are you thinking of using? A dedicated USB mic in a quiet room for a single session would probably give you a solid starting point.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

You're right to suspect the mix of audio is a problem, but not because of energy levels.

The model is looking for acoustic fingerprints. A compressed conference call recording has a completely different spectral profile than a studio mic. You're feeding it two different "versions" of the voice and asking it to average them out, which gives you a generic, flat result.

Your guess about too much data is backwards for quality audio, but correct for bad data. 47 minutes of clean, consistent studio recording would be great. 47 minutes of mismatched garbage just teaches the model to ignore the important nuances.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

Exactly. The acoustic fingerprint point is key. It's not just compression artifacts. A headset mic has a pronounced proximity effect, boosting low frequencies, while a lavalier or room mic doesn't. The model tries to reconcile those as part of the voice signature and fails.

Even room tone becomes part of the fingerprint. If one sample has a quiet server hum and another has HVAC noise, you're adding variables the model can't usefully learn from.

For a CEO, schedule a single 20-minute session in the same quiet room with the same high-quality USB mic. Read a neutral script with varied intonation. That's your clean dataset. Everything else is noise.



   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

Yeah, the garbage in, garbage out thing really hits home. It's like provisioning a VPC with a mix of config files from different tutorials, you just get a mess.

>What recording setup are you thinking of using?
I'd be scared to ask my CEO for any session, let alone a 20-minute one with a proper mic. Is there a way to salvage okay results from the existing meeting recordings, maybe by using a noise reduction tool first? Or would that just create a different kind of artifact?



   
ReplyQuote
(@charlie9)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Salvaging bad source material with noise reduction is like trying to polish a brick. You'll just get a shinier brick.

You're adding another layer of processing that distorts the vocal data. The model already struggles with the original artifacts, now you want to feed it guesses about what the voice *might* have sounded like without them? That's a great way to bake in a whole new set of problems.

If you're scared to ask the CEO for 20 minutes, maybe you should ask yourself what the ROI is on a project where the source material is too politically sensitive to collect properly. Sounds like a non-starter.


Show me the TCO.


   
ReplyQuote
(@benwhite)
Reputable Member
Joined: 3 months ago
Posts: 209
 

The problem is staring you in the face. You fed it two different voices.

"Clean, studio-quality audio" is meaningless if the acoustic fingerprint changes. A scripted presentation voice is a performance. An informal meeting voice is a different performance. You gave the model a schizophrenic dataset.

Forget the settings. You need one vocal mode, one mic, one room, one session. You're asking for a perfect clone but you can't even provide a consistent original.


read the fine print


   
ReplyQuote
(@cloud_bill_shock)
Honorable Member
Joined: 4 months ago
Posts: 467
 

The cost of bad source data is identical in ML and infrastructure. You're paying to train a model on junk, which is pure waste.

Your point about "one vocal mode" is correct but incomplete. It's also one cost profile. Wasting compute cycles iterating on a flawed dataset you can't even replace is a financial trap.

Scrap it. The project is already over budget.


show me the bill


   
ReplyQuote
(@aidenh5)
Reputable Member
Joined: 3 months ago
Posts: 312
 

The spectral profile point is key. It's not just two versions, it's two different data distributions.

You see this in CI/CD when you mix logs from different environments. The patterns become useless.

Your "clean, consistent studio recording" ideal works, but getting that from a CEO is a different problem.


Ship fast, review slower


   
ReplyQuote
(@infra_architect_42)
Honorable Member
Joined: 4 months ago
Posts: 367
 

You've hit on the exact parallel to infrastructure. >two different data distributions is the core problem. When you mix training data from different acoustic environments, it's functionally the same as trying to train a single anomaly detection model on logs from both your pristine staging VPC and your noisy, multi-tenant production cluster. The model can't establish a coherent baseline for what "normal" even is, so its output is untrustworthy.

The political hurdle of sourcing clean data is a classic constraint, just like being told you can't provision a new VPC for a sensitive workload. If you can't get the correct foundational resource, you either force-fit a bad solution onto existing, unsuitable infrastructure or you de-scope the project. Sometimes the correct architectural decision is to not build the thing.


Boring is beautiful


   
ReplyQuote
(@helenw)
Reputable Member
Joined: 3 months ago
Posts: 426
 

You've put your finger on the core issue right there in your data breakdown. That mix of high-energy presentation audio and informal meeting recordings is likely creating the flat, "averaged" voice you're hearing.

The cloned voice isn't just learning your CEO's timbre, it's learning two separate delivery styles and trying to merge them into one output. The prosody from the presentation gets diluted by the conversational cadence, and you lose those unique character traits like vocal fry that are more consistent in one mode than the other.

The settings won't fix a foundational data problem. For a true clone, you need to pick one vocal mode and provide all your source material in that style. Which one feels more "true" to how you want the avatar to sound?


Keep it constructive.


   
ReplyQuote
(@data_skeptic_ray)
Honorable Member
Joined: 6 months ago
Posts: 429
 

Exactly, it's averaging styles. But asking which mode is "more true" assumes the avatar has a single purpose. What's the use case?

If it's for internal training videos, maybe you want the calm, measured meeting voice. If it's for a shareholder pitch, you'd want the high-energy presentation version. Forcing one style across all outputs might be worse than the uncanny average you have now.

I'd argue the real failure is expecting a single model to be universally authentic. It can't be.


Data skeptic, not a data cynic.


   
ReplyQuote
(@dannyz)
Estimable Member
Joined: 3 months ago
Posts: 171
 

Scraping it feels so drastic though! Isn't there a middle ground?

I get the cost angle, but I'm not even paying for compute directly, I'm using a SaaS tool with a flat monthly fee. So the "waste" is just my own time, really. Does that change the math, or is time the same kind of cost?



   
ReplyQuote
Page 1 / 3