Skip to content
Notifications
Clear all

Has anyone tried the 'whisper' mode? Does it actually sound like a whisper?

44 Posts
42 Users
0 Reactions
27 Views
(@cloud_infra_rookie)
Noble Member
Joined: 4 months ago
Posts: 552
Topic starter   [#28112]

Hi everyone! I’m still pretty new to TTS and voice cloning tools, so I hope this isn’t a silly question.

I was looking at Resemble AI’s features and saw they have a ‘whisper’ mode. Has anyone here actually tried it? I’m curious if the output genuinely sounds like a real, breathy whisper, or if it’s more like a quiet normal voice. I’m thinking of using it for a short audio drama project, but I want it to sound convincing. Any examples or tips would be awesome! 😊



   
Quote
(@alexh42)
Reputable Member
Joined: 3 months ago
Posts: 227
 

Good question, not silly at all. I've tested it on a project.

It's definitely more than just a quiet voice - you get that characteristic breathy texture. The main caveat is it can sound a bit "clean" or processed compared to a real human whisper, which has more mouth noise and inconsistency. For your audio drama, I'd recommend generating a few versions and maybe layering in a very subtle ambient room tone to ground it. It works best for short phrases, not long paragraphs.



   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

It's not a silly question, it's exactly the kind of practical detail that matters.

For audio drama, you'll probably need to do some post-processing to sell it. The synthesized whisper often lacks the slight lip smacks, throat clearness, or uneven breath control of a real one. Think of the AI output as a clean vocal track - you'll want to add a separate, very subtle layer of those organic mouth sounds underneath it in your editor.

If you're generating longer dialogue, break it into shorter phrases. The model can struggle with maintaining that whisper consistency over extended sentences, sometimes slipping back into a more normal vocal register mid-way.


Build once, deploy everywhere


   
ReplyQuote
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
 

It's convincing enough for a first pass, but you'll run into the same issue as any synthesized voice: it lacks entropy.

A real whisper has random breath peaks and drops in amplitude. This mode gives you a consistent, flat breath layer. For drama, you'll hear the difference.

Process the output with a compressor set to a high ratio and fast attack. That will mimic the uneven push of air. And as others said, keep the input phrases short.


Trust but verify, then don't trust.


   
ReplyQuote
(@chrisf)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Oh, the entropy point is really interesting, hadn't thought of it that way. Makes total sense.

The compressor trick is a great idea. I'm still new to audio editing - would any basic compressor plugin in something like Audacity do the trick, or do you need more specific controls for that?


Still learning.


   
ReplyQuote
(@gracep)
Reputable Member
Joined: 3 months ago
Posts: 297
 

Yes, I've run it through some spectral analysis.

It's a synthesized whisper. The core issue is it lacks the high-frequency air turbulence (the "shhh" noise) that dominates a real whisper's spectrogram. You're getting vocal fold vibration with a breathy layer, not actual fricative noise from teeth and tongue.

For audio drama, you'll hear the difference. The compressor trick mentioned helps with dynamics, but you're missing that key spectral component.

Layer a very low-volume white noise or pink noise bed under the track, high-pass filtered above 1kHz. Adjust the level until it just blends. That gets you closer to the missing texture.


Data over opinions


   
ReplyQuote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

No, it's not a silly question. You need measurable validation for audio projects.

The other replies are correct about the missing spectral texture and entropy. The compressor and noise layer advice will help, but they're post-processing fixes for a core synthesis limitation.

My caveat: if your audio drama involves headphones or close-mic listening, the synthetic cleanliness will be more noticeable. For general speaker playback, it might pass. Always test your final mix on the intended output device.


Five nines? Prove it.


   
ReplyQuote
(@benchmark_hunter)
Reputable Member
Joined: 6 months ago
Posts: 341
 

Good question. The output is convincingly breathy, but it suffers from what I call "benchmark voice syndrome": too perfect. It lacks the subtle vocal fry and random breath catches that make a human whisper feel intimate. If you're using it for drama, run a few generations with slightly different temperature settings - that can introduce just enough variability to break up the synthetic consistency.


Numbers don't lie


   
ReplyQuote
(@first_timer_evan)
Reputable Member
Joined: 4 months ago
Posts: 278
 

The "benchmark voice syndrome" is a great way to put it. That synthetic perfection is exactly what I'm trying to avoid when looking for a natural sound.

When you say to adjust the temperature settings for variability, does that usually mean a higher setting? And have you found that it introduces more of those breath catches, or does it sometimes just make the pronunciation weird?



   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

Good point about the high-frequency fricatives being absent. That's the clinical giveaway for me. Your noise layer fix is smart.

I'd add a small caveat: if the whispered dialogue is meant to be extremely intimate, like an ASMR style, the pink noise trick can sometimes feel like a separate soundbed and break the illusion. It works best when the voice is mixed with other ambient sound or music. For a solo, clean whisper, the synthetic core might still poke through.


—AF


   
ReplyQuote
(@danielh)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Yep, "benchmark voice syndrome" nails it. That perfect consistency is the dead giveaway.

For temperature, I've found a *slightly* higher setting (like 0.8-0.9) can help, but it's a gamble. It sometimes introduces a bit of welcome rasp, but other times just makes the cadence weird. The real trick is to generate several short phrases at different temps and stitch the good bits together in post.

Makes you appreciate the random little imperfections in a real human voice, doesn't it?


Keep deploying!


   
ReplyQuote
(@garethp)
Estimable Member
Joined: 3 months ago
Posts: 226
 

The multi-generation approach you described is the only way I've found to inject useful unpredictability. The downside is it's essentially brute-forcing entropy, which becomes expensive at scale.

You're right about the gamble - I've seen higher temperature produce clipped sibilants that sound like digital artifacts rather than breath catches. If the model architecture isn't training on whispered data specifically, you're just introducing general instability.

For anyone stitching clips, watch for subtle shifts in the room tone or noise floor between generations. You'll often need a light noise gate or consistent low-cut filter across all segments to hide the seams.


Plan the exit before entry.


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 4 months ago
Posts: 668
 

Yeah, the temperature gamble is real. > higher setting? can introduce more of those breath catches

In my tests, a higher setting *can* sometimes add a bit of that vocal fry texture, but it's not reliably generating actual breath catches. More often, it just makes the pacing feel uneven or slurs consonants. It's like turning up the randomness on a perfectly smooth synth pad - you get weird artifacts before you get natural-sounding imperfections.

For what it's worth, I've had better luck generating a few takes at the default temp and then manually adding in subtle breath sounds from a separate library. It's more work, but you avoid the pronunciation roulette.


cost first, then scale


   
ReplyQuote
(@data_diver_42)
Honorable Member
Joined: 7 months ago
Posts: 400
 

The separate breath library trick is solid for a one-off. Makes me wonder about the economics though - at what point does stitching a sound library onto every clip become more expensive than generating 5-10x the takes and cherry-picking? Feels like a manual ETL job vs. a brute-force query.

Have you found any libraries where the breath textures actually match the timbre of the synthesized voice? That's always my hangup, the breaths sound like they're from a different person.


Data is the new oil - but it's usually crude.


   
ReplyQuote
(@amandak9)
Reputable Member
Joined: 3 months ago
Posts: 209
 

Exactly the trade-off I've been weighing on my current project. That manual stitching *does* feel like ETL, and the cost curve flips if you're generating more than a few dozen clips.

> libraries where the breath textures actually match the timbre

It's a constant battle. I've had decent results by processing the breath sample *through* the same final vocal chain as the whisper - a touch of the same EQ, reverb, and subtle distortion. It glues them together a bit, but it's still an imperfect match. Sometimes I think the only real solution is to record a single breath from the target speaker and clone it.


Show me the accuracy numbers.


   
ReplyQuote
Page 1 / 3