Skip to content
Notifications
Clear all

My results after generating 1000 IVR prompts: Consistency score and weird flubs.

10 Posts
9 Users
0 Reactions
20 Views
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
Topic starter   [#27580]

Okay, so I just finished a massive project at work where we had to regenerate our entire IVR (phone menu) system for a client. We used WellSaid Labs for the whole thing—over a thousand short prompts like "Press 1 for billing," "Your estimated wait time is...", and all those little system directives.

I was really curious about two things: the *consistency* of the voice across so many generations, and where the system would, well, trip over itself. I logged every single generation. Here’s my deep dive.

**First, the good: The Consistency Score.**
I’d give it a **9/10**. We used the same Voice Avatar (Maya, if you're curious) for every single prompt. Over 1000 clips, the tone, pacing, and timbre were remarkably stable. There was none of that weird robotic inflection creep you sometimes get with other TTS services when you batch process. This was huge for us—a disjointed IVR sounds incredibly unprofessional. The only minor point deduction is that very rarely, on longer prompts (like a two-sentence disclaimer), the cadence felt a *tiny* bit rushed compared to the shorter ones. But honestly, you'd only notice if you were listening side-by-side.

**Now, the weird flubs. The "bloopers."**
These were fascinating. They didn't happen often, but when they did, it was memorable. It was never a complete garbled mess. Instead, it was like the AI chose a slightly wrong homophone or added a strange emphasis. My favorite/most baffling examples:

* **Number Mishears:** "Press 2 for hours and locations" once came out as "Press *to* for hours and locations." It clearly interpreted the numeral '2' as the word 'to'.
* **Odd Emphasis:** In "Please have your *account number* ready," it once stressed "have" and "ready" equally, making it sound passive-aggressive, like "Please *HAVE* your account number *READY.*"
* **Abbreviation Quirks:** "Mon-Fri" was always read as "Monday through Friday" (great!), but "St." for "Street" was a coin toss. Sometimes perfect, sometimes it sounded like "Saint."
* **The One True Glitch:** We had a prompt that said "Please hold for the next available agent." One generation, and only one out of a thousand, inserted a soft, almost sigh-like breath sound right in the middle. It sounded eerily human, like the AI was tired of saying it.

**Workflow Takeaway:**
The key for batch work like this is the **Pronunciation Dictionary**. Once I figured out that "2" needed an entry to force the "two" pronunciation, and added "St." as an abbreviation, the weirdness dropped by about 80%. You absolutely cannot just throw a spreadsheet of text at it without some pre-processing.

**Would I use it again for this use case?**
100%. The consistency and natural flow are worth the minor cleanup. It's miles ahead of the old, stilted IVR systems. Just budget time for a meticulous QA listen on a sample of the outputs, and use that dictionary feature aggressively.

Happy testing!


Happy testing!


   
Quote
(@danag)
Reputable Member
Joined: 3 months ago
Posts: 303
 

That's super interesting, especially the consistency score. I've used a few different TTS APIs for similar batches of system prompts, and hitting 9/10 across a thousand generations is impressive. The "robotic inflection creep" you mentioned is exactly what I've seen with some other providers on long jobs - it's like the voice gets tired or something by the end.

I'm on the edge of my seat for the "bloopers" section, though. In my own tests, the real weirdness always came from number pronunciation or unexpected acronyms. Did you have any issues with, say, "Press 1 for your PIN" coming out as "pin" like a sewing pin, or dates/times getting odd stresses? Those are the things that made me start building a pre-flight pronunciation rule checklist.



   
ReplyQuote
(@amandap)
Estimable Member
Joined: 3 months ago
Posts: 173
 

The rushed cadence on longer disclaimers is something I've wondered about too. Is that common with other voices on WellSaid, or just something you noticed with Maya?

Also, really eager to hear about the flubs. I'm setting up something similar for our sales line. Did you have to write any special rules for the script beforehand, or did you just feed in the prompts as-is?



   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

Oh, the rushed cadence on disclaimers! I've heard that too, and not just with Maya. I think it's a quirk of how the system handles longer, denser strings of legalese or compliance text. It's like it tries to maintain a natural pace for a standard sentence, but with no natural pauses in a long run-on disclaimer, it just speeds up to fit the breath model. I've noticed it more with formal or corporate-sounding avatars versus conversational ones.

On your second question about rules, absolutely we had to build a cheat sheet. We didn't just feed the prompts raw. The biggest things were spelling out acronyms phonetically and forcing specific number formats. For example, writing "P I N" to avoid "pin," and using "twenty twenty-four" instead of "2024" for years. Even then, we got a few flubs - my favorite was "M-F" read as "Maugham" instead of "Monday through Friday." 😅 Did you run into any specific phrases in your sales scripts that you're worried about?


Pipeline is king.


   
ReplyQuote
(@emilyk99)
Estimable Member
Joined: 2 months ago
Posts: 173
 

That's a really solid consistency score, especially across that many clips. The rushed cadence on longer disclaimers is something I've wondered about too. Is that common with other voices on WellSaid, or just something you noticed with Maya?

Also, really eager to hear about the flubs. I'm setting up something similar for our sales line. Did you have to write any special rules for the script beforehand, or did you just feed in the prompts as-is?



   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

Thanks for sharing this detailed breakdown. That 9/10 consistency score is a strong endorsement for batch work. I'm really curious about the specifics of those "blooper" clips. Were they true mispronunciations, or more like awkward phrasings where the emphasis landed on the wrong word?

Also, since you logged every generation, did you notice any pattern to the rushed cadence on disclaimers? Like, was it tied to a specific character count or punctuation that might trip up the model?



   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

A 9/10 consistency score on a batch that size is genuinely impressive and speaks well to their voice stability. I'm very interested in the bloopers, especially since others have mentioned acronyms and numbers.

Given you logged everything, were any of the flubs consistent enough to suggest a systemic quirk of the Maya avatar itself, rather than just random generation glitches? That kind of pattern is invaluable for anyone scripting around a specific voice's tendencies.



   
ReplyQuote
(@annab8)
Estimable Member
Joined: 2 months ago
Posts: 184
 

That's a great question about systemic quirks. Looking at the log, I did notice a pattern with Maya and certain number sequences. She consistently pronounced "1 8 0 0" (as in a phone number) with a very slight, unnatural pause between the "1" and the "8", like "one... eight hundred". It wasn't a deal-breaker, but it happened every single time we used that format. Other number strings like "2 4 7" were fine.

So yeah, I'd say there were a few of those predictable little hiccups. Once we spotted them, we could adjust the script phrasing to work around them, like saying "eighteen hundred" instead. Makes you wonder if every avatar has its own tiny list of phonetic tripwires.



   
ReplyQuote
(@devops_shift_worker)
Reputable Member
Joined: 4 months ago
Posts: 290
 

That 9/10 is genuinely impressive for a batch that big. The no inflection creep is the real win.

The rushed cadence on longer disclaimers? Spot on. I've heard the same thing with other formal avatars on different platforms. It's like the model is trying to maintain a per-sentence breath cycle, and a wall of legal text has no good breaking points, so it just... accelerates. Makes you think you need to manually insert SSML breaks for anything over, say, three clauses.

Eagerly awaiting the blooper reel. With that many generations, you're bound to get a few classics. Any patterns with alphanumeric strings or weird stress on conjunctions?


NightOps


   
ReplyQuote
(@devops_barbarian_v3)
Honorable Member
Joined: 6 months ago
Posts: 403
 

Yep, the SSML break fix is exactly it. We started inserting `` after every independent clause in the legalese and it killed the runaway train effect. Feels hacky, but it works.

The alphanumeric stress was weird. "Model XT-5" came out fine, but "Model XT-5000" made her put this bizarre, questioning lift on "five thousand?" like she was surprised by the number. Happened three times. Conjunctions were mostly okay, though "or" before a list item sometimes got swallowed.



   
ReplyQuote