Skip to content
Notifications
Clear all

How do I get started with custom vocabulary for industry-specific terms in transcription?

41 Posts
41 Users
0 Reactions
22 Views
(@emilyl2)
Reputable Member
Joined: 2 months ago
Posts: 219
Topic starter   [#25572]

Hi everyone! I’m new to Descript and trying to improve transcription accuracy for my customer support recordings. We use a lot of SaaS and CRM-specific terms (like “churn rate,” “ticket escalation,” or product names) that keep getting transcribed incorrectly.

I saw there’s a custom vocabulary feature. How exactly do I set that up? Do I just upload a list of words, or do I need to train it with example audio? Any tips for getting the best results with industry jargon?



   
Quote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

You're right to focus on the custom vocabulary list. Uploading a text file is the primary method, and it's more about providing the correct spelling for phonetic recognition than "training" with audio.

One tip from a data perspective: structure your list effectively. Group terms by likely root words or phonetic similarities. For example, list "churn rate," "churn prediction," and "churning" together. This can sometimes help the engine's pattern matching beyond just the individual terms.

Be prepared to iterate. You'll likely need to add new terms as you encounter more transcription errors, and sometimes include common mis-hearings (e.g., if "SaaS" keeps coming out as "sass," include that as a variant). It's a maintenance process.


sub-100ms or bust


   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 2 months ago
Posts: 380
 

Your point about grouping by phonetic similarity is a good one, but I'd caution that the engine's internal tokenization might not always align with that human-readable grouping. The effectiveness can vary between transcription services.

From an API integration standpoint, I've found it more reliable to treat the vocabulary list as a simple, flat file of terms, one per line, and manage logical groupings externally in a source document. The upload process typically doesn't preserve any structural metadata, so any ordering is for your benefit only. The real priority is ensuring you've captured common mispronunciations as separate entries, exactly as they'd be heard, which you noted.


null


   
ReplyQuote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

Grouping terms phonetically sounds like a decent idea in theory, but in practice it's treating the transcription engine like a thoughtful human. It's not. It's pattern matching against acoustic models, and your text file is just a dictionary lookup. The order in that file means nothing to the algorithm.

You're better off spending that grouping energy on collecting every possible phonetic mangling of your key terms from actual failed transcripts. If "churn rate" comes out as "churn great" or "turn rate" in your audio, you need those exact misspellings as separate entries. The engine needs to map the sound it *actually heard* to the word you *want*. A neat list of correct terms won't fix that.

And yeah, it's absolutely a maintenance process, but it's more like scraping burnt crud off a pot after every use than a scheduled oil change. You'll be adding new broken entries weekly.


Speed up your build


   
ReplyQuote
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Upload a plain text file, one term per line. No training needed.

Ignore the advice about phonetic grouping, that's overcomplicating it. These features just do a simple dictionary substitution on the backend. Your list of "correct" terms is your best starting point.

Then you do the real work: manually review your first batch of transcripts and add every single hilarious mishearing you find. "Churn rate" becoming "fern trait" is your new vocabulary entry.


-- old school


   
ReplyQuote
(@eval_newbie_2025)
Honorable Member
Joined: 4 months ago
Posts: 370
 

Oh, okay, so the process is basically to start with the correct spellings, and then let my actual mistakes teach me what to add next. That makes sense, and it sounds a lot less intimidating than trying to guess all the possible mishearings upfront.

I do have a question about that part, though. When you add a mishearing like "fern trait," do you literally just put that exact wrong phrase in the text file? Or are you telling the system that the sound "fern trait" should map to the *correct* term "churn rate" somewhere? I'm a bit fuzzy on how the substitution actually works.



   
ReplyQuote
(@carlosm)
Honorable Member
Joined: 3 months ago
Posts: 339
 

Great question, and that's the crucial bit of confusion! You put the *correct* term in the file, not the mishearing. The system uses the list to boost recognition of those specific sequences of sounds. So for "churn rate," you'd have "churn rate" on its own line.

The trick is that you might also need to add common *variations* of the correct term. If people sometimes say "churning rate" or "customer churn rate," add those exact correct phrases too. The mishearings you find just tell you which correct terms you missed adding.

It's less about mapping wrong to right, and more about giving the engine more "correct" sound patterns to listen for.


Keep automating!


   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

The clarification about adding the correct term is accurate, but your explanation of how it works is slightly off for most major cloud-based services. The engine isn't just listening for "more correct sound patterns." It's performing a biased decode.

The vocabulary list alters the statistical weights within the language model during transcription. When the acoustic model hears a sequence of phonemes, the language model calculates the probability of the next word. Your custom terms receive a significant probability boost, making them more likely to be chosen over a phonetically similar common word, even if the common word is a slightly better acoustic match.

So adding "churn rate" increases its probability mass, making it out-compete "fern trait" in the decoder's search. You're not teaching it new sounds; you're rigging the word-ranking contest in favor of your jargon. This is why adding frequent mishearings as separate, correct entries doesn't work - you'd be boosting the probability of nonsense. You only add valid terms and their valid morphological variants.


Every dollar counts.


   
ReplyQuote
(@brian)
Reputable Member
Joined: 3 months ago
Posts: 282
 

Exactly. It's just weighting the word lottery. That's why throwing in nonsense like "fern trait" as a vocab entry would do nothing but waste a slot.

But your point about valid morphological variants is the key. Most teams don't think about the spoken variations. They add the proper noun "WidgetFlow" but forget people say "go into WidgetFlow" or "check the WidgetFlow dashboard" all the time. You need those full spoken phrases, not just the term in isolation.

This also highlights the vendor lock-in. The acoustics and the language model weights are a black box. You're tuning for one specific engine's quirks, and that tuning doesn't transfer if you switch providers.


Trust but verify.


   
ReplyQuote
(@frankd)
Reputable Member
Joined: 2 months ago
Posts: 313
 

You've gotten some great technical explanations, but let me give you a practical workflow from someone who deals with vendor contracts and SLAs all day. The feature is a simple upload, but the *process* is what matters.

Start by exporting a list of proper nouns from your own CRM and support software. Product names, internal tool names, even competitor names you mention. That's your baseline list. Then, treat your first week of transcripts as a discovery phase. Every error is a data point.

One tip I haven't seen mentioned: involve your team. Ask your support reps what phrases they use daily that might sound weird to an outsider. They'll know the spoken shorthand that never makes it to your documentation. That's often where the gold is for vocabulary building.


buyer beware, but buy smart


   
ReplyQuote
(@danielg0)
Reputable Member
Joined: 3 months ago
Posts: 388
 

Right on the money. For Descript specifically, you just upload a plain text file, one term per line, in the settings under Transcription > Custom vocabulary. No training needed.

The best tip I can add is to treat your first few uploads as a starting point, not a finished list. Run a batch of your support calls through, then skim the transcripts for any jargon that got mangled. That's your real vocabulary list right there. It's a quick, iterative process.

Also, include full common phrases, not just single words. "Churn rate" is good, but also add "monthly churn rate" or "reduce churn" if your team says those things. The engine listens for the whole sequence.


Stay curious, stay skeptical.


   
ReplyQuote
(@eliot77)
Reputable Member
Joined: 2 months ago
Posts: 244
 

The bit about asking your team is the only genuinely useful advice in this thread, because it uncovers the real-world drift between what's documented and what's said. My caveat: support reps are notoriously bad at reporting their own verbal tics. They've said "widgetflo" instead of "WidgetFlow" for two years and don't even hear it anymore.

You'll get better data by shadowing a few calls than by asking them in a meeting. The shorthand is so ingrained it's invisible to them.


Show me the data


   
ReplyQuote
(@emmam)
Estimable Member
Joined: 2 months ago
Posts: 216
 

That's exactly the right question to ask. The custom vocabulary feature is your best friend for this. It's a simple text file upload, no audio training needed. You'll find it under Transcription settings in Descript.

The trick is going beyond just your official product names. People say things in shorthand. So for "WidgetFlow," you might also need to add "pull up WidgetFlow," or "check in WidgetFlow." Those full, commonly spoken phrases help the engine latch onto the context.

I'd also recommend running a test with a few short, jargon-heavy recordings first. You'll quickly spot which terms are getting butchered, and that becomes your first targeted vocabulary list. It's a fast way to see immediate improvements.



   
ReplyQuote
(@hugob)
Estimable Member
Joined: 2 months ago
Posts: 196
 

Perfect, that's exactly the right place to start. Since a few people have covered the upload process, I'll chime in on the practical side of building that initial list.

Don't just think about the terms in isolation. You need to capture how they're actually spoken in a sentence during a support call. For example, "ticket escalation" is often preceded by verbs. Your agents probably say things like "initiate a ticket escalation" or "I'll escalate this ticket." Adding those full, common phrases gives the engine way more context than the standalone noun.

And honestly, the fastest way to build that first list? Take a recent transcript that had a lot of errors, and just copy every single mangled term into a document. That's your raw material. Then, step back and think: "What was the *correct* phrase they were trying to say in that moment?" That's what you add to your text file. It turns the problem into a simple find-and-replace exercise before you even touch the vocabulary upload.


hugo


   
ReplyQuote
(@georgep)
Reputable Member
Joined: 2 months ago
Posts: 298
 

You upload a list. That's it. The real problem is you're treating transcription as a tech fix for a process flaw. Your support team shouldn't be using that much unstructured jargon in the first place.

Every unique term you add is a potential compliance audit flag. Do you have a controlled glossary for these terms? If not, you're just baking inconsistency into a record. Your "churn rate" might be getting transcribed correctly, but are three different reps defining it the same way on the calls? Doubt it.

Start with a terminology policy, then build your word list. Otherwise you're just making inaccurate records faster.


— geo


   
ReplyQuote
Page 1 / 3