Skip to content
Notifications
Clear all

Help: custom vocabulary not being respected for product names.

2 Posts
2 Users
0 Reactions
0 Views
(@davidn)
Estimable Member
Joined: 4 days ago
Posts: 56
Topic starter   [#17385]

I’ve been conducting a structured evaluation of MeetGeek over the past three weeks for our B2B logistics team. A key requirement is accurate transcription of our internal product codes and client-specific terminology.

I configured a custom vocabulary list with approximately 50 entries, primarily multi-word product names and model numbers (e.g., "ThermoSafe Shipper 48L", "RFID-PALLET-2024"). The goal is to force the transcription engine to treat these as single, unbroken terms.

However, in my test meetings, the vocabulary is not being consistently respected. The system often:
* Inserts spaces where there shouldn't be any ("Thermo Safe Shipper").
* Splits alphanumeric codes ("RFID PALLET 2024").
* Sometimes ignores the terms entirely, reverting to its best phonetic guess.

I’ve followed the documented process: CSV upload, confirmed the list is active for the correct meeting types, and ensured the language setting (en-US) matches.

Has anyone else implemented a similar technical vocabulary for specialized inventory or product nomenclature? I’m looking to isolate the variable:

* Is there a character limit or format constraint for vocabulary entries that isn't documented?
* Does the engine prioritize phonetic interpretation over the provided dictionary in certain conditions?
* Has a successful workaround been found, such as spelling the terms phonetically in the vocabulary list itself?

My next step is to run controlled tests with single-word versus multi-word entries, but community insight would help direct that analysis.


Measure twice, buy once.


   
Quote
(@danielr23)
Trusted Member
Joined: 1 week ago
Posts: 67
 

I hit the same issue with manufacturing part numbers. The vocabulary list is more of a suggestion than a rule. It works best on isolated words, not alphanumeric strings or compound names.

Check your raw audio quality first. If there's any background noise or crosstalk, the model will default to its own phonetics. A separate test with a clean recording of just the terms can isolate if it's an audio problem versus a vocabulary processing problem.

Also, their documentation is wrong about the CSV format. You need to wrap entries containing hyphens or spaces in double quotes, even though they say it's optional.


Trust, but verify


   
ReplyQuote