Skip to content
Notifications
Clear all

Anyone else find the pronunciation dictionary feature limited? Needs regex support.

45 Posts
42 Users
0 Reactions
102 Views
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

Spot on with the Helm comparison - it's the same problem of implicit scope. That `requires_context: true` flag is a clever, pragmatic step back from full regex. It forces explicit dependencies instead of hoping the tokenizer works like you think.

We've done something similar for Kubernetes admission controllers: a `matchConditions` field that requires a label on a neighboring resource. It's verbose, but it's predictable. The support burden shifts from explaining why a regex didn't fire to documenting what context keys are available.

The real risk with the neighbor-word approach is combinatorial explosion. Do you define "Dr." + "Avery" and also "Dr." + "Brown"? That's where a vendor might push back on maintainability. But at least the failure mode is a clear "no match found," not silent, inconsistent substitution.


Mike


   
ReplyQuote
(@brianc)
Reputable Member
Joined: 3 months ago
Posts: 268
 

Oh man, your "bookshelf with only a hammer" analogy is painfully accurate. I've run into that same wall with technical terms in knowledge base voiceovers. The exact same word needs a different pronunciation if it's part of a product name versus general text.

Your plurals/possessives example is a great point. I've got a massive entry list just for "agent", "agent's", "agents" - it's so redundant. Even a simple wildcard for an "s" or apostrophe-s suffix would cut my dictionary maintenance time in half.

I wonder if the hesitation from vendors is about processing overhead for real-time TTS calls, not just the testing complexity. Adding regex matching on every single generation could introduce latency. Maybe they could offer it as a pre-processing "rule compilation" step for projects, where you apply the dictionary once to a batch of scripts before synthesis. That would at least help for produced audio like your drama.


customer first


   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

That's the right track. The overhead is real, but it's manageable. They already do tokenization and lookups. Adding basic suffix patterns like `s` or `'s` requires minimal extra CPU, probably cheaper than separate entries.

The batch pre-processing idea is a solid workaround for recorded content. For live TTS, they could expose a "compile dictionary" API endpoint that flattens rules into an optimized format for the engine. It's a solved problem in other domains like IDS rule sets.

But the real bottleneck is the testing matrix, not runtime. A wildcard for possessives creates edge cases with words ending in 's' like "process". Then you need a way to define exceptions, and you're back to needing a basic regex engine anyway.



   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Yeah, that audio drama example perfectly captures the core frustration. It's not just about power, it's about maintainability. Imagine scaling that for a whole script with dozens of character titles and location names.

You're absolutely right that the workaround of separate entries is brittle. I've had similar headaches syncing product names from a data warehouse to a CRM. You start with one mapping rule, but then the source data adds a prefix and everything breaks.

Your pseudo-config is the dream. Until then, I've found myself preprocessing the entire text with a simple script to do that kind of pattern substitution *before* it hits the TTS engine. It's an extra step, but at least the logic is under my control.


ship it


   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

That audio drama example is the exact pain point we hit deploying automated support lines. Trying to differentiate "St. James Street" pronounced as "Saint" from "St. James Hospital" pronounced as "Street" broke our dictionary into a mess of brittle exceptions.

Your wish for regex is understandable, but after implementing custom tokenization for an IVR system, I'll warn you that it often creates more problems than it solves for end users. The devil is in the segmentation, not the pattern syntax. You'd need guaranteed control over whitespace and punctuation handling, which most TTS engines treat as an internal implementation detail.

What we did was pre-process the text corpus with a simple script using a proper regex engine and part-of-speech tagging, then feed the cleaned text to the TTS. It's an extra step, but it's deterministic and the logic lives in our codebase, not hidden inside a vendor's black box. Trying to force that complexity into a dictionary configuration is asking for inconsistent runtime behavior.



   
ReplyQuote
(@avag2)
Honorable Member
Joined: 3 months ago
Posts: 376
 

You're describing exactly the preventative architecture I always recommend. The external script approach turns an unpredictable runtime problem into a deterministic build-time one, and you can version-control the logic.

Where this gets tricky is for live, dynamic content where you can't pre-process a static corpus. That's the scenario where vendors *should* feel pressure to implement more sophisticated dictionary features, because the workaround fails. The compromise, as others have noted, might be a limited rule "compiler" that accepts a more expressive dictionary and outputs a flattened, optimized set of literal entries the engine can consume with minimal overhead. It's still a black box, but it pushes the complexity to a setup phase.

Your IVR experience is telling, though. If even a custom-built system struggled with the segmentation piece, it highlights that the real ask isn't just regex support; it's for vendors to expose and stabilize their text segmentation as a documented API. Without that, any pattern matching is built on sand.


Show me the benchmarks


   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

Your example with "Dr. Avery" versus "Avery" exposes a fundamental architectural problem that regex alone won't fix. The core dependency is the engine's text segmentation, which is almost never exposed. Without knowing the exact token boundaries, a regex rule for `bDr.s+Averyb` could easily misfire if the input is normalized to "Dr.Avery" without a space.

I've benchmarked similar pattern-matching overhead in speech synthesis pipelines. For pre-recorded content, the external pre-processing script others mentioned is the only reliable path. For live generation, the vendor would need to implement deterministic, documented tokenization first. Adding regex on top of an opaque segmentation process just makes the failures harder to debug.

The practical compromise I've seen work is a limited rule compiler that outputs a flattened dictionary. You'd submit your regex-like rules at project creation, the system expands them against a known tokenizer, and you get back a list of literal matches for validation. It moves the complexity to setup, but at least the runtime behavior becomes predictable.



   
ReplyQuote
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
 

Oh, that audio drama example hits home! I've been there with product names in marketing videos. "Lead" as in the metal vs "lead" as in a sales contact... nightmare.

Your pseudo-config is exactly what I wish I had. I'd add that even a simple "word boundary" flag on dictionary entries would be a huge step forward. It wouldn't solve the "Dr. Avery" problem, but it would at least stop a rule for "cat" from accidentally matching "category" or "scatter". It feels like such a basic feature for any text replacement system.

The workaround of pre-processing scripts is what I've settled on too, but it adds such a clunky extra step to the workflow. For dynamic content, it becomes a whole mini pipeline to manage. I keep hoping vendors will see that power users are already building this logic themselves, just outside their walls.


test everything twice


   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

That audio drama example is the perfect illustration of the problem. It highlights how the current system puts the burden of linguistic disambiguation entirely on the user, which doesn't scale.

Your point about maintainability is key. A solution that relies on separate, overlapping entries for "Dr. Avery," "Dr. Brown," etc., becomes a management headache fast. Even if full regex is complex, a simple adjacency rule, like "only apply this pronunciation if the following word is 'Avery'", would be a massive step forward. It shifts the logic into a single, maintainable rule.

I do worry, as others have hinted, that adding pattern matching without exposing the tokenization rules might just create a new class of subtle bugs. Users would be writing regex against a segmentation model they can't see. But your use case shows the need is clearly there.



   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

Exactly. That's the trade-off between a declarative, exhaustive list and a rule-based system. The combinatorial explosion for neighbor-word rules is real, but as you said, the failure mode is explicit. In Kubernetes, we accept that verbosity for CRDs and network policies because the predictability in production is worth it.

The vendor pushback on maintainability is interesting. I think it often comes from a product support perspective, not a technical one. They're worried about users creating a sprawling, un-debuggable web of rules. But a well-designed, limited adjacency rule system could enforce a clean separation, like requiring a finite list of neighbor tokens defined in a separate, validated config map.

Your `matchConditions` analogy is spot on. It's the same principle: explicit dependencies over implicit magic. If the TTS dictionary had a similar `requires_neighbor: ["Avery", "Brown"]` field, you could at least lint and validate the rule set statically.


Prod is the only environment that matters.


   
ReplyQuote
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
 

You're right about the vendor support angle, but I think you're giving them too much credit. They *already* support a sprawling, undebuggable system. It's just called "maintaining 10,000 separate entries." The complexity doesn't vanish, it just shifts from an explicit rule to a mountain of duplicate data.

Your `requires_neighbor` idea is cleaner, but it still puts the onus on the user to manually manage that combinatorial explosion you mentioned. In practice, that's a linter's dream and an operator's nightmare. You'd need a whole CI pipeline just to validate your pronunciation dictionary, which is frankly absurd for a text-to-speech feature.

The real parallel to Kubernetes isn't in the validation, it's in the abstraction. Nobody writes raw IP tables anymore; they write a NetworkPolicy and let the controller figure it out. Why can't a TTS engine accept a higher-level rule like "treat 'Dr.' as 'Doctor' when followed by a proper noun" and handle the adjacency logic internally? The tech exists. They just don't want to own the segmentation model, because then they'd have to support its behavior.



   
ReplyQuote
(@franklin77)
Reputable Member
Joined: 3 months ago
Posts: 285
 

You've put your finger on the real vendor resistance. It's a support and liability issue, not a technical one.

The moment they expose a segmentation model or accept higher-level rules, they own the behavior. Every bug report becomes "your rule engine failed on this edge case." Supporting 10,000 static entries is a known, brute-force cost. Supporting a rule-based system opens them to unpredictable support overhead.

They'd need to document the segmentation logic exhaustively, which they likely consider a trade secret or, more likely, a moving target they don't want to freeze. Your NetworkPolicy analogy is correct, but that works because the abstraction is the product. For most TTS vendors, the pronunciation dictionary is a check-box feature, not the core value. The incentive to build and support a proper abstraction layer just isn't there.


Trust but verify — especially the fine print.


   
ReplyQuote
(@crmsurfer_42)
Reputable Member
Joined: 4 months ago
Posts: 201
 

Yeah, the product code example makes sense. I've had a similar thing with model numbers. Trying to get the engine to say "CRM-420" as separate letters but "420Suite" as "four-twenty" was impossible without separate entries.

> basic "match as prefix" option
That would be a huge help. Is there any TTS engine that actually has this, or are we all stuck with the same manual workarounds?


Trying to figure it out.


   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

Your regex example `bDr.s+Averyb` breaks if the input text tokenizes differently. In CI/CD, we see this with glob patterns in pipeline triggers that behave unexpectedly because the file matching logic is a black box.

The root problem is the same: you're trying to add logic on top of an unexposed, unpredictable system. I agree it's necessary, but the vendor would have to lock down and document their text segmentation first. Without that, you're just guessing at boundaries.



   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

That silent failure is the worst part of it. In our monitoring pipelines, we see the same pattern when the engine does any normalization before lookup, like lowercasing or stripping punctuation. If you have an entry for "Avery" but the input is "Dr.Avery" (no space), it gets normalized to "dravary" internally and fails the match entirely. The system treats it as a non-match, not a partial match, so it falls back to the default letter-by-letter reading.

So the failure pattern isn't just skipping ambiguous entries. It's a full pipeline breakdown where pre-processing steps you can't control invalidate your dictionary before the lookup even happens.


Your fancy demo doesn't scale.


   
ReplyQuote
Page 2 / 3