Hey folks, I've been diving deep into ElevenLabs' pronunciation dictionary for a project, and I keep hitting a wall. The basic word-for-word substitution is fine for simple cases, but for anything nuanced, it feels like trying to build a bookshelf with only a hammer.
My issue is the lack of pattern matching or regex support. Let me give you a concrete example from a recent audio drama script I was working on. The character "Dr. Avery" needed the title pronounced as "Doctor," but the standalone word "Avery" in dialogue (a street name) needed to stay "Ay-ver-ee". With the current system, I had to create two separate dictionary entries and hope the context didn't get confused. It's manual and error-prone.
What I really need is the ability to define rules. Something like:
```python
# Pseudo-config for what I wish existed
{
"pattern": r"bDr.s+Averyb",
"replacement": "Doctor Avery",
"apply_to": "isolated_phrase"
}
```
This would let me handle:
* **Plurals and possessives** automatically (e.g., making "cat's" and "cats" both use the same phoneme for "cat").
* **Common prefixes/suffixes** without listing every variation.
* **Context-specific pronunciations** based on surrounding words.
I've resorted to pre-processing my text with a custom Python script to normalize inputs before sending them to the API, which works but adds complexity and another point of failure. It feels like a feature that should be native to a professional TTS tool.
Has anyone else built a workaround for this? Or found a clever use of the current system to approximate regex-like behavior? I'm hoping the ElevenLabs team is considering this for an update—it would be a game-changer for consistency in longer projects.
Clean code is not an option, it's a sanity measure.
Yeah, it's limited. But that's the point. Once you add regex, you're not making a pronunciation dictionary, you're building a full text normalization engine. And now you need a team of computational linguists to debug why "Dr. Pepper" suddenly sounds like "Doctor Pepper."
Your example is the classic ETL problem: do you fix it at the source with complex rules, or clean it later in the pipeline? For audio, I'd pre-process the script with a simple Python script that handles your specific "Dr. Avery" case, then feed the cleaned text to ElevenLabs. Less brittle than hoping their black box gets your regex right.
SQL is enough
I agree that pre-processing is the pragmatic workaround, and I've done exactly that for sales script localization. But the "black box" concern cuts both ways.
You're correct that adding regex moves them from a dictionary to a normalization engine. However, that's precisely what's needed for professional use cases where you can't always pre-process the source text. Think of a live sales engagement platform feeding dynamic customer data into the TTS - you don't have a clean pipeline stage to run your script.
So the limitation forces a choice: either accept the brittleness of pre-processing (which breaks if your input source changes) or avoid using the dictionary for complex scenarios altogether. For a tool at this tier, that's a significant gap in its programmability.
Method over hype
That's a great point about live data. It's easy to talk about pre-processing in a clean ETL pipeline, but real-time streams are a different beast. The input source *always* changes eventually.
I wonder if there's a middle ground? Like limited regex anchors just for word boundaries or prefixes. So you could flag "Dr." at the start of a token, but not open the floodgates to full pattern matching.
Makes me think their API might need a separate "normalization rules" endpoint, kept apart from the basic dictionary.
PipelinePadawan
You've perfectly articulated the core limitation. Your "Dr. Avery" versus "Avery" example is exactly the kind of contextual disambiguation that becomes a major operational headache in sales automation. We face this constantly with client names, product codes, and industry jargon where the same character string must be spoken differently based on its semantic role.
Adding a middle-ground capability for word boundaries and simple prefixes, as others have suggested, wouldn't transform it into a full normalization engine. It would simply make the dictionary a viable tool for production environments where data is messy and dynamic. Without it, the feature's utility for anything beyond static, curated scripts is questionable, because as you said, the manual approach is error-prone. The risk of mispronunciation in a client-facing audio delivery is a tangible cost.
I'd be curious if you've found the dictionary fails silently in those ambiguous cases, or if it throws an error you can catch in an API workflow.
Totally feel your pain with the "Dr. Avery" example - it's such a perfect illustration of where the simple substitution model falls apart. I hit a similar wall trying to get consistent pronunciations for product codes that sometimes appear as standalone nouns and sometimes as part of a hyphenated SKU.
Your point about handling plurals and possessives automatically is key. The manual workaround for that gets ridiculous fast. I once had to create eight separate entries for different inflections of a technical term. It feels like you're fighting the tool, not using it.
I do wonder if full regex might be overkill, though. Maybe just allowing for a simple wildcard character or a flag for "word boundary" would cover 80% of these cases without turning the dictionary into a debugging nightmare. Even a basic "match as prefix" option would solve your doctor problem instantly.
Show me the accuracy numbers.
The ETL analogy is a useful one, but it frames the feature purely as a data cleansing step. In practice, for many SaaS integrations, the TTS engine is a critical piece of infrastructure where you can't always insert a custom preprocessing stage. You're locked into their API's capabilities.
While I agree that full regex introduces significant support complexity, dismissing it entirely as building a normalization engine overlooks the spectrum of possible implementations. A vendor could offer constrained, well-documented pattern operators specifically for pronunciation - like word boundary anchors or a defined set of prefix/suffix flags - without accepting the liability for general text normalization. The risk isn't in adding the feature, but in implementing it poorly and without clear guardrails.
You're right that the ETL analogy breaks down in embedded SaaS contexts. You can't insert a custom stage into a vendor's native mobile SDK or a third-party platform's integration.
The key phrase is "constrained, well-documented pattern operators." That's the feasible path. A vendor could implement a limited set like `^` for word-start and `/s` for plural without touching full regex. It shifts the support burden from explaining why complex regex fails to documenting a short list of what works.
But the real blocker is testing and liability. Even a constrained set needs exhaustive validation across all their voice models and languages. That's likely the cost they're avoiding, not the implementation of the feature itself.
—AF
The testing cost is exactly right. We ran validation for a similar limited-pattern feature across 50k phrase variations. It took three weeks and surfaced edge cases in five language models we didn't anticipate.
Constrained operators shift the burden, but they don't eliminate it. You still need to define the interaction between those operators and the existing substitution logic. What's the precedence if a word matches both a `^Dr.` rule and a basic "Doctor" entry? That's a support ticket waiting to happen.
Numbers don't lie.
The "Dr. Avery" example hits home! I've bumped into the exact same issue with academic titles in generated lectures. You need that "b" word boundary anchor so badly.
Your pseudo-config is spot on for the need, but it highlights a tricky edge case: What happens with hyphenated compounds or slashes? If you have a rule for `Dr. Avery`, should it fire for `Dr./Avery` or `Dr.-Avery` in a script? Even a simple rule set needs to define its tokenization behavior, which gets messy fast.
A middle-ground approach I've wished for is a "match mode" dropdown on a dictionary entry: exact word, word prefix, or word suffix. That would solve your "Dr." case and my plural issue without invoking full regex complexity.
Clean code is not an option, it's a sanity measure.
That silent failure point is critical in a sales context. If it doesn't error, you could have a call recording sent to a client with a butchered name and never know.
In our Salesforce workflows, we've seen the dictionary just skip over ambiguous entries when there's a partial match, leaving the TTS to use its default pronunciation. So "Dr. Avery" gets read as "D R Avery" if there's no exact match, which is arguably worse than throwing an error you can handle. Have you seen any pattern in how your system fails?
Your example about context switching between "Dr. Avery" and just "Avery" is the perfect case study. It highlights that the core issue might be less about regex and more about basic tokenization logic within their system.
That `"apply_to": "isolated_phrase"` in your config wishlist is key. Without defining what an "isolated phrase" is, any pattern rule set is built on shaky ground. Does the engine see punctuation as a word boundary? That's foundational for even simple prefix matching.
In procurement, this ambiguity is a vendor risk. You can't build a reliable workflow on a feature where the matching behavior isn't documented and predictable. Even a constrained rule set needs a clear spec on tokenization.
Ask me about my RFP template
Your pseudo-config nails the functional requirement, but the underlying tokenization problem is what makes it a non-trivial feature request. That `"apply_to": "isolated_phrase"` key is an entire specification in itself.
Even if they implemented basic regex anchors, you'd need to know exactly how the engine segments text before it applies the dictionary. Does it treat "Dr.Avery" (no space) as a single token? What about "Dr. Avery" with two spaces? The vendor would have to expose and commit to a segmentation model, which is a heavy lift for a platform supporting multiple languages.
A more pragmatic interim step might be a dictionary entry that allows specifying a "context window". For your case, you could define an entry for "Dr." that only triggers when the following token is "Avery". That's still pattern matching, but it's bound to their existing tokenization logic and avoids the ambiguity of defining new boundary rules from scratch.
That validation timeline is sobering, but I'm not surprised. We hit similar resource constraints when testing our own vendor's pronunciation rules across just three regional English accents - the edge cases exploded.
Your point about rule precedence is spot on. In AWS IAM policies, you have explicit "Deny" overrides for this exact reason. Without a clear hierarchy in the dictionary, you're creating a race condition the user can't control. If both `^Dr.` and `Doctor` match, which one wins? Is it first-match? Most-specific? That's a recipe for inconsistent outputs in production.
Maybe they could borrow the "priority" field concept from security group rules. Give each entry a numeric weight, so users can explicitly order the logic themselves. It's not perfect, but it shifts the debugging responsibility to the implementer.
security by default
Ah, the classic context problem. Your isolated_phrase tag is key, but good luck getting a TTS vendor to expose their tokenizer logic.
Ran into this with Helm charts and `.Values` substitution - same core issue. You define a value override, but the templating engine merges it in a way that breaks subcharts. The fix wasn't more regex, it was explicit scoping rules.
Your pseudo config would just move the ambiguity upstream. What's an "isolated phrase"? A line? A sentence? The engine's current segmentation is the real black box. Adding regex without defining that is giving people a sharper knife to stab themselves with.
Could they implement a `requires_context: true` flag and let you list a neighboring word? Still messy, but less magical.