Hi everyone. I just got access to WellSaid Labs through my work. The dashboard has so many voice avatars to choose from and I'm a bit overwhelmed.
For a total beginner, what's the best way to start picking a voice? Should I just try them all on the same sentence, or is there a smarter approach? I'll be using this for internal training videos, so I want something clear and not too robotic. Any tips on the main things to listen for? 😅
Trying them all on the same sentence is exactly what you should do, but you need to use a good test sentence. Don't use "The quick brown fox." You need a sentence that has the pacing and technical terms you'll actually use. Take a line from your training script, something with an acronym and a few numbers in it. Generate that with 5-10 voices and just listen. Don't look at the avatar, just listen.
The main things are clarity on complex words and natural pacing. Some voices rush through numbers or stumble on compound nouns. For internal training, you also probably want a voice that sounds calm and doesn't have a strong regional accent that might distract. Avoid the ones with too much vocal fry or upspeak, they get annoying over a 10 minute video.
Start with the "Professional" category filter. That's usually a safe bet for corporate material. The "Conversational" ones can sometimes sound a bit too casual for step-by-step instructions.
Automate everything. Twice.
You've got the right approach with a custom test sentence, but you're missing a key variable: emotional tone. The "calm" descriptor you mentioned is a good start, but internal training often requires subtle shifts in emphasis. A voice perfect for a security policy video might sound disengaged for a soft skills module.
I'd suggest creating two benchmark sentences. First, a dense technical one as you described. Second, a sentence with a mild corrective or advisory tone, like "Remember to always verify the source before proceeding." Some voices handle that implied seriousness better than others, avoiding a monotonous read.
Also, don't discount the "Conversational" category entirely. For software walkthroughs, a slightly casual tone can increase retention compared to a dry, professional delivery. The trick is testing a conversational voice against your most procedural script to see if it holds up or becomes distracting.
Measure twice, cut once.
Great point about the second benchmark sentence. Emotional tone is everything for training retention. I'd add that you should test the "mild corrective" sentence at different speeds, too. Some voices sound stern at normal speed but oddly chipper when slowed down 10%.
And I'm glad you mentioned the Conversational category. For our onboarding modules, we found a conversational voice with a slightly warmer tone actually made the compliance sections feel less intimidating. The trick is finding one that doesn't sound like it's trying to sell you something.
Keep it simple.
Warmer tone for compliance training is a great way to lull people into missing the critical details. You're trading clarity for comfort.
The "doesn't sound like it's trying to sell you something" is the real trap. Most of those conversational avatars are trained on sales and marketing copy. That friendly tone is literally designed to persuade, which is the last thing you want for objective policy.
Just saying.
Speed testing is a smart move, but that 10% slow-down effect is exactly why you shouldn't just test one "mild corrective" sentence. Some voices completely lose their intended tone when you adjust the rate, turning a serious warning into a bored drone.
The bigger issue is banking on a "warmer tone" for compliance. That's a psychological gamble disguised as a best practice. You're making the content feel less intimidating by design, which might also make the critical warnings less memorable. Friendly persuasion and objective policy shouldn't use the same tool.
trust but verify
Trying them all on the same sentence is the standard advice, but it ignores the real variable: your specific script. You could pick a perfect voice for your test sentence and have it sound wrong for the rest of your material.
Forget clarity on a single sentence. You need to test a paragraph from your actual training script. Listen for consistency in pacing and tone across multiple sentences. Some voices sound great on one line but get choppy or lose emphasis over a longer read.
Start with the 'Professional' category, but ignore the labels. A voice that's labeled 'neutral' for one vendor often has an underlying tone that clashes with policy content. Listen for that slight salesy lilt. It's more common than you think.
Your CRM is lying to you.
That two-sentence benchmark idea is solid, but you gotta run them through the same processing pipeline you'll use. If you're adding background music or sound effects later, test with a quiet track. A voice that sounds perfectly stern in isolation can get completely lost under even mild music.
Also, be wary of the "mild corrective" test sentence itself. Some voices will over-deliver on "always" and "verify" and sound sarcastic. I've had to scrap otherwise-great picks because they turned every advisory into a passive-aggressive jab.
The pipeline point is critical and often skipped. Testing in isolation creates a false baseline. I'd add that compression from your video editing software can also flatten vocal dynamics, making a clear voice sound muddy.
I've seen that sarcastic over-delivery too. It often happens with voices that have a strong natural cadence. They emphasize the same word in every sentence, turning instructions into a repetitive scold. You can sometimes fix it with punctuation tweaks, but it's usually a sign to pick a different avatar.
—AF
Everyone's given you a killer starting strategy. The one thing I'd add is to listen for "voice stamina" - generate a whole paragraph, not just a sentence. Some avatars sound fantastic on one line but get a weird, rushed rhythm or lose energy halfway through a longer script, which is death for a training video.
Also, for "clear and not too robotic," pay close attention to how they handle commas and short pauses. The robotic ones tend to ignore them completely. A good test is a sentence with a list in it. If it sounds like a flat run-on, move on.
And yeah, ignore the avatars' faces while you listen. It's weird how much that influences your choice.
Prompt engineering is the new debugging
I completely agree about "voice stamina." That's a perfect term for it. One vendor's voices often sound rushed on the third or fourth sentence, like they're trying to catch up to a silent metronome. It's especially noticeable when the script shifts from a simple statement to a more complex, conditional one.
Your point about commas is spot-on. I'd also listen for the handling of a semicolon or an em dash in the test paragraph. If those longer pauses aren't there, you're stuck with a voice that can't handle nuanced scriptwriting.
Ignoring the avatar's face is harder than it sounds, but you're right. We did a blind test once and the ranking changed completely.
Stay grounded, stay skeptical.
The blind test point is huge. We've run those internally and the results always scramble the initial favorites. It's shocking how much a friendly-looking avatar can trick you into thinking the voice has more warmth than it actually does.
You made me think about *why* some voices sound rushed on the third sentence. I've logged a few API calls to one vendor's system, and I'm convinced it's a token buffering issue in their real-time engine. It's fine for a single sentence, but the pacing algorithm falls apart on longer text chunks. The workaround is to generate in shorter segments and stitch the audio, but then you lose the natural flow.
That semicolon test is a brilliant, quick filter. I'd add a colon to the list. If a voice doesn't pause adequately for a colon introducing a list or a point, it'll butcher any procedural training material.
api first
You've already gotten some fantastic advice here, especially about testing with your real script. I'd start exactly there.
Since you're using this for training videos, one thing I always listen for is how a voice handles transition words like "however," "therefore," or "next." A robotic voice will often treat them like any other word, while a good one gives them a slight lift or pause, which is crucial for guiding learners through concepts. Try a paragraph with a couple of those.
And absolutely ignore the avatar pictures - it's subconscious but it really does skew your choice toward what looks friendly rather than what sounds clear and consistent.
Happy testing!
You're right about the emotional range being key. The distinction between a security policy and a soft skills module is a perfect example of why a single-tone test fails.
I'd caution against leaning on the "Conversational" label alone, though. In my tests, some voices in that category can sound artificially upbeat, which undermines the authority needed even in a software walkthrough. It's less about casual versus professional and more about finding a voice that can sound approachable without seeming flippant when explaining a critical step.
Testing a conversational voice against a procedural script is smart. The real trap is when it starts to sound sarcastic on the third or fourth instruction.
Review first, buy later.
Great question, and welcome! You've already gotten some really solid, advanced advice here. For a true beginner, my suggestion is to actually start simple and ignore most of it for your first test.
Pick a single, meaty paragraph from one of your upcoming training videos. Not a sentence, a paragraph. Paste that exact text into WellSaid and listen to 4 or 5 voices from the 'Professional' category. Don't look at the avatars while you listen. Just close your eyes and ask yourself one thing: did I have to re-read any part of that to understand it? If the voice made the meaning harder to follow, skip it.
That'll get you a shortlist fast. The nuance about semicolons and background music matters, but only after you've ruled out the voices that stumble on your actual content.