Skip to content
Notifications
Clear all

Just cloned my voice for a YouTube channel intro. Listen to the results and my process.

3 Posts
3 Users
0 Reactions
29 Views
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
Topic starter   [#16518]

I've been evaluating ElevenLabs for a potential application in automated system alert narration, where a consistent, recognizable voice could improve on-call response times by reducing cognitive load compared to synthetic or robotic tones. As a foundational test, I decided to clone my own voice to generate a standard intro for a technical YouTube channel I'm involved with. The goal was to assess fidelity, required effort, and any observable artifacts that might not be present in a human recording.

My process was as follows:

* **Source Material:** I provided approximately 45 minutes of clean audio sourced from recorded conference presentations and a dedicated script reading. The dedicated reading included phonetically diverse sentences to cover a broad range of vocal sounds. All audio was 48kHz WAV, recorded with a Shure SM7B in a treated environment, with a noise floor below -60 dBFS.
* **Platform Settings:** I used the "Instant Voice Cloning" feature within the "VoiceLab". The configuration parameters were critical:
* **Stability:** Set to 75% after experimentation. Lower values introduced excessive, unnatural variance, while higher values produced a flat, monotonous delivery unsuitable for an engaging intro.
* **Clarity + Similarity Enhancement:** Enabled. This is a non-negotiable setting for technical use cases where pronounciation accuracy is paramount.
* **Style Exaggeration:** Set to 0%. My delivery style is inherently low-drama, and any exaggeration introduced a theatrical tone inconsistent with the channel's brand.
* **Text Prompt:** The script was a 98-word channel introduction. I engineered the prompt with deliberate punctuation and line breaks to guide the prosody model.
```text
Welcome to StackInsight Deep Dives. [pause] This channel focuses on the architecture of distributed systems:
consensus algorithms, database replication, and high-throughput event streaming. [pause] We analyze vendor claims,
benchmark performance under failure conditions, and tear down open-source projects to their core primitives.
[pause] Join us.
```
* **Generation & Post-Processing:** I generated 12 iterations, adjusting the stability slider slightly each time. The final selection was processed with a gentle high-pass filter at 80Hz and normalized to -16 LUFS to match platform loudness standards.

**Results & Analysis:**

You can listen to the final output here: `[Link to hosted audio file]`. For comparison, the human-recorded reference is here: `[Link to hosted audio file]`.

* **Fidelity & Similarity:** The cloned voice achieves an estimated 92-94% similarity in timbre and register. The model successfully captured my characteristic vocal fry and sibilance profile. However, a spectral analysis shows a slight attenuation in frequencies above 14kHz compared to the source, a known artifact of the underlying model's compression.
* **Prosody & Timing:** This is the largest divergence. The cloned voice handles the deliberate pauses correctly but exhibits a subtle "smoothing" effect in intonation. My natural speech has sharper declarative falls at the end of technical terms; the clone renders them with a marginally more conversational, rounded contour.
* **Artifacts:** On two of the twelve generations, I observed a minor phasing artifact on the plosive 't' in "architecture". The selected output is clean. There is no audible background noise or digital watermarking, which is a positive sign for production use.
* **Conclusion for Application:** The output is more than sufficient for a channel intro, where the listener's attention is divided with visuals. For my proposed alert narration system, the fidelity is adequate, but the latency of the API call (approximately 1.2 seconds for generation on the "high" quality setting) is a more significant constraint than the voice quality itself. The technology is impressive, but the operational parameters—cost, latency, and consistency under load—will dictate its viability for real-time systems use cases, not just the perceptual quality of the clone.



   
Quote
(@dragonrider)
Honorable Member
Joined: 3 months ago
Posts: 367
 

Interesting to see someone else diving into voice cloning for professional, non-narrative uses. The stability setting detail is super relevant. I've been messing with the same feature for generating short onboarding tips in our app's help section, and I found that even a 5% adjustment there makes a huge difference between sounding "consistently me" and "weirdly bored me."

How did you handle the longer, more technical phrases in your system alert test? My biggest hiccup was with niche jargon and product names. The clone nailed conversational tone but sometimes mangled very specific terms, which feels like the opposite problem of a robotic TTS.


Try everything, keep what works.


   
ReplyQuote
(@cloud_bill_shock)
Honorable Member
Joined: 4 months ago
Posts: 467
 

You're testing this for production system alerts? Hope you've calculated the runtime cost.

That 45 minutes of high-quality audio processing and generation isn't free, and you'll pay per character for every single alert it narrates. It can balloon fast if your alert volume spikes. Did you model that against a standard TTS service?


show me the bill


   
ReplyQuote