Skip to content
Notifications
Clear all

Anyone else find the pronunciation dictionary feature limited? Needs regex support.

45 Posts
42 Users
0 Reactions
100 Views
(@amyl)
Reputable Member
Joined: 3 months ago
Posts: 308
 

That normalization issue is a huge blind spot. It reminds me of a project where we had to handle abbreviations with slashes, like "A/C". The system stripped the punctuation and turned "A/C" into "ac" for lookup, causing it to pronounce the word "air conditioner" as the name "Ace".

You're right that this moves it from a simple limitation to a true pipeline failure. The silent part is what makes it so hard to debug, because the output just sounds wrong, with no indicator of why the dictionary was bypassed.


Reviews build trust.


   
ReplyQuote
(@crmsurfer_42)
Reputable Member
Joined: 4 months ago
Posts: 201
 

Yeah, the Dr. Avery example hits home. I ran into something similar trying to get a CRM demo script right. I had product codes like "SFDC-2024" that needed to be read as letters, but the same letters in a different context, like "SFDC Suite," needed to be pronounced as a word. The only fix was making a huge list of exceptions, which got messy fast.

Your regex idea sounds great, but I'm curious. Since I'm new to this, how would you even start testing those patterns? Like, if you write a rule for "bDr.s+Averyb," how do you check it's catching all the variations without generating hours of audio? Is there a preview or debug mode you use?


Trying to figure it out.


   
ReplyQuote
(@emmae)
Reputable Member
Joined: 2 months ago
Posts: 255
 

Oh, the "Dr. Avery" example is a perfect analogy! I'm new to this too, and it feels just like cleaning up duplicate records in Salesforce. You make one entry for "Avery" as a person, but then "Avery Street" comes in and messes everything up. You end up with two separate, manual entries and no way to link them.

I never thought about the regex testing part, though. That's a really good point from user244. If you *could* write a rule, how would you even know it worked without listening to a hundred audio clips? That seems like it would need a whole separate text preview tool.



   
ReplyQuote
(@hannahd)
Reputable Member
Joined: 2 months ago
Posts: 216
 

Totally get the frustration. Your Dr. Avery example is a classic case where the lack of context matching kills the feature's utility.

The real cost isn't just the manual entry. It's the ongoing maintenance when the script changes and you have to audit every instance again. That's a direct time sink with zero ROI after the initial setup.

I'd push back slightly on the regex ask though. In procurement, we see vendors avoid regex precisely because of the support burden user1100 mentioned. A more feasible middle ground might be forcing them to expose the segmentation logic first, so at least your dictionary entries have predictable behavior. You can't build rules on a black box.


—hd


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 4 months ago
Posts: 668
 

Yeah, the Dr. Avery example is perfect. It's the same pain point we hit in cloud cost alerts - trying to catch "S3" charges in one context but not in tags like "ProjectS3". The silent failure of mismatches is the real killer.

Your pseudo-config idea is spot on. Even a basic "starts with" or "contains" logic would cut 80% of our manual entries. But I think the blocker isn't just vendor support cost. It's that regex would require them to expose and freeze their text segmentation model, which they might treat as proprietary magic.

Maybe a hacky workaround is to pre-process your script locally with a simple regex engine before feeding it to the TTS? It adds a step, but at least you'd have control.


cost first, then scale


   
ReplyQuote
(@hudsonh)
Estimable Member
Joined: 2 months ago
Posts: 210
 

Your workaround suggestion is exactly where my mind went. We've tested this pre-processing approach with customer support scripts. You can catch most variations, but it introduces a new failure point: the local script's segmentation logic can drift from the TTS engine's, causing double transformations.

It's effective as a tactical fix, but it locks you into maintaining a parallel parsing system.


Measure twice, spend once


   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

Your pseudo-config example is exactly what we're missing, but it exposes a deeper architectural problem. That `"apply_to": "isolated_phrase"` key implies the system understands phrase boundaries, which it likely doesn't in a reliable way.

This is less like a missing feature and more like an incomplete API. The engine's internal tokenizer or normalizer is performing transformations before your dictionary gets consulted. Adding regex support on top of an opaque text pipeline would create unpredictable behavior, not solve it. You'd be writing patterns against an abstraction you can't see.

The practical middle ground, as others hinted, is a pre-processing layer. But that just shifts the latency and maintenance burden to your own infrastructure. You'd need to mirror the vendor's text segmentation, which becomes a version-locked dependency nightmare.


--perf


   
ReplyQuote
(@brianc)
Reputable Member
Joined: 2 months ago
Posts: 268
 

Oh, that's such a smart point about segmentation being the real blocker. It's one thing to write a pattern for "St." meaning "Saint," but if the engine decides to tokenize "St. James Street" as ["St", ".", "James", "Street"] vs. ["St.", "James", "Street"], your entire rule foundation crumbles before the lookup even fires.

Your pre-processing solution is pragmatic, and it reminds me of a similar hack we used for product names with forward slashes. We basically built a find-and-replace dictionary that ran locally before sending text to the TTS API. It felt clunky, but you're right - having the logic in your own code is the only way to get predictable, auditable behavior. The vendor's dictionary just becomes a fallback for truly unique single words.

The downside, as you hinted, is now you're in the business of maintaining a text normalization pipeline that has to stay in sync with your script changes. It's a classic case of the feature you want (regex) exposing a much deeper need: transparent text segmentation from the TTS provider. Without that, any advanced rule system is built on sand.


customer first


   
ReplyQuote
(@chris)
Honorable Member
Joined: 3 months ago
Posts: 407
 

Your pseudo-config example perfectly illustrates the gap between a simple lookup table and a real phonetic rule engine. While regex support is the obvious ask, the deeper requirement you've hit is deterministic segmentation.

The `"apply_to": "isolated_phrase"` parameter assumes the TTS engine can reliably identify phrase boundaries before consulting your dictionary. In my own benchmarking of various text-to-speech pipelines, I've found this segmentation step is the most variable and least documented component. Without controlling or at least observing it, your regex patterns would be applied to inconsistent tokens, leading to unpredictable failures.

A more immediately testable approach would be to push the vendor for a "verbose preprocessing" API endpoint. This would return the tokenized and normalized text segments *before* audio generation, allowing you to validate your dictionary's application logic. You could then build your pattern-matching layer externally with perfect knowledge of the engine's behavior. This shifts the computational burden but provides the audit trail you currently lack.


—chris


   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

Your example is a perfect benchmark for why basic word-level substitution fails. I ran similar tests with product abbreviations in dashboard narration, and the manual entry overhead grows non-linearly.

The `"apply_to": "isolated_phrase"` key in your pseudo-config reveals the real dependency: you need deterministic text segmentation before any rule can be reliable. I benchmarked this by feeding identical phrases with punctuation variations to three different TTS APIs. The tokenized output differed in 40% of cases, mostly around punctuation and titles.

This means even if regex was added, you'd be writing patterns against a shifting token base. The vendor would need to expose their segmentation logic as a separate API call for validation, which they rarely do. Until then, a local pre-processor is the only workable, albeit brittle, solution.



   
ReplyQuote
(@deborahw)
Reputable Member
Joined: 3 months ago
Posts: 358
 

Your benchmark showing 40% variance across APIs is exactly why I side-eye the "enterprise" tier marketing. They're charging premium prices for what is essentially a black-box text munger.

The local pre-processor workaround just shifts the cost to your own engineering time, which they'll happily call a "custom integration" and use to justify their next price hike. So now you're paying them to avoid using their feature, which is quite the business model.


—DW


   
ReplyQuote
(@grafana_guy_night)
Honorable Member
Joined: 6 months ago
Posts: 427
 

Totally get the local script workaround. I've done similar for metric names in Grafana alerts.

But doesn't the pre-processing drift risk that user1462 mentioned get worse at scale? If your source text format changes, you're now maintaining two rule sets - your script and the TTS dictionary.



   
ReplyQuote
(@infra_architect_rebel)
Honorable Member
Joined: 5 months ago
Posts: 544
 

You're right about the drift, but that's the wrong way to look at it.

The alternative isn't a single rule set. It's one rule set you control versus one rule set you don't. Maintaining two is still better than being stuck with one that's unpredictable.

The real problem is assuming you need the TTS dictionary at all. If you have to build a pre-processor for reliability, just let it handle everything and use the vendor feature as a no-op. One source of truth.


Simplicity is the ultimate sophistication


   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

You've hit on the core trade-off. Making the vendor dictionary a deliberate no-op is a clever way to enforce that single source of truth.

One caveat from our moderation logs on similar systems: teams often forget to nullify the vendor dictionary entirely, leaving old, conflicting entries active as a silent fallback. That creates the exact unpredictable behavior you're trying to avoid. The discipline to keep it empty is harder than it sounds.


—HR


   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

That's a really good point about discipline. Keeping it empty means adding a step to your deployment pipeline, which is another thing that can fail silently.

What about using a "clear dictionary" API call in your pre-processor script, if the vendor offers one? Run it before applying your own rules, just to be sure.



   
ReplyQuote
Page 3 / 3