I've been evaluating WellSaid Labs for generating technical narration and training materials. The voice quality is impressive for standard prose, but we hit a consistent snag when our scripts include proper nouns (like "Kubernetes," "Istio," "Prometheus") or specific jargon (e.g., "gRPC," "Envoy," "Terraform"). The default pronunciation often butchers these terms, breaking the flow.
I understand custom phonemes are the official solution, but managing a separate phoneme dictionary for a dynamic set of technical terms across multiple projects and voices seems like a scaling headache. It feels like we'd be building and maintaining a parallel infrastructure just for pronunciation.
Has anyone developed a workflow or pattern to handle this more efficiently? I'm thinking along the lines of:
* **Pre-processing scripts:** Lightly modifying SSML input texts to spell out tricky words phonetically (e.g., "Kubernetes" -> "Koo-ber-nay-tees") before sending to the API.
* **A shared lexicon:** A simple, version-controlled JSON file mapping terms to their phonetic approximations, used by a pre-render step.
* **Caching strategy:** For frequently used terms, pre-rendering short clips and stitching them into the final audio, though this complicates editing.
The goal is to keep the pipeline as GitOps-friendly as possible and avoid manual tweaks in the WellSaid Studio UI. What's been your experience? Are we better off just embracing the custom phonemes feature and building tooling around it?
The shared lexicon is a decent stopgap, but you're just reinventing their custom phonemes with less support. Why maintain your own shadow infrastructure when the vendor's solution is the actual product feature?
Your pre-processing script idea creates a new problem. Who maintains the mapping? How do you validate the phonetic approximations actually sound right across different voices? You'll spend more time debugging "Koo-ber-nay-tees" than if you'd just built the phoneme list once.
Have you actually tested the scaling headache of custom phonemes, or are you assuming it's bad? A single, well-structured dictionary attached at the project level isn't a parallel infrastructure. It's a config file.
Caveat emptor.
You're describing exactly what we already do at my shop. We have a single, project-level phoneme dictionary. It's not a parallel infrastructure, it's a config file. We version it alongside our IaC definitions.
The scaling headache isn't in maintaining the dictionary, it's in not having one. You'll waste more time post-processing audio clips and debugging approximations than you would just defining "Kubernetes" once.
Proof in production.
Totally feel your pain on this. My team deals with the same with marketing automation platforms (try getting "Marketo" or "Eloqua" said correctly out of the box 😅).
I think your pre-processing script idea has legs, especially for a dynamic set of terms. We used a similar approach as a temporary bridge. The key caveat we found was that phonetic approximations are voice-dependent - what sounds decent in one voice can be way off in another. You'll want to bake in a voice-specific mapping layer.
Honestly, after a few months, we bit the bullet and built a centralized phoneme dictionary anyway. The pre-processor just became the maintenance layer for *that*. It's less of a scaling headache than I feared, mostly because the list of truly problematic jargon plateaus. New terms pop up, but it's manageable.
Maybe try the script for a sprint and see how many mappings you actually create? It'll tell you pretty quickly if the dictionary is the simpler path.