Skip to content
Notifications
Clear all

Migrated from Otter to MeetGeek 6 month report on accuracy and reliability

38 Posts
38 Users
0 Reactions
122 Views
(@hudsonh)
Estimable Member
Joined: 2 months ago
Posts: 210
Topic starter   [#24592]

After six months of using MeetGeek as our primary meeting transcription and analysis tool, replacing Otter.ai, I have enough data to share a substantive review. The migration was driven by our need for better integration with our CRM and more granular conversation analytics for sales coaching. This report focuses on the core metrics of accuracy and reliability, which are foundational for any team building processes around these transcripts.

**Accuracy Benchmarks (Internal Sales Calls):**
We conducted a bi-weekly spot-check on 10-minute segments across 50 different calls, comparing MeetGeek's output to manual transcriptions. The environment includes a mix of clear audio and typical remote meeting challenges (background noise, crosstalk).

* **Overall Word Accuracy:** Averaged **94.2%** over the period. This is a marginal improvement over our final Otter benchmark of 92.8%. The difference is most noticeable in industry-specific jargon.
* **Speaker Diarization Accuracy:** MeetGeek correctly identified and separated speakers **87%** of the time in meetings with 3+ participants. This is a notable strength. Otter often struggled here, especially when voices were similar.
* **Critical Error Rate:** We tracked "critical errors"β€”misheard words that change the meaning of a sentence (e.g., "won't" vs. "want," pricing numbers). MeetGeek averaged 1.2 per 30-minute call, compared to Otter's 1.8.

**Reliability & Workflow Observations:**
* **Recording Failures:** We experienced two complete meeting recording failures in six months (both traced back to user error with calendar permissions). This is comparable to our Otter experience.
* **Processing Time:** The average time from meeting end to finalized transcript delivery was **8 minutes**. This increased to ~15 minutes for calls over 60 minutes. Performance has been consistent month-to-month.
* **Integration Downtime:** No observed downtime where transcripts failed to push to our Salesforce environment via Zapier. The automated workflow has proven more stable than Otter's native integrations, which occasionally required re-authentication.

**The Pitfall: Keyword Highlighting vs. Context**
While not strictly an accuracy issue, MeetGeek's keyword highlighting for "action items" or "questions" can be misleading. It reliably surfaces sentences with question marks but often misses declarative action items ("I'll send that by Friday"). This required custom training for our team to treat these highlights as starting points, not definitive summaries.

For teams where transcript data feeds into coaching or compliance, the improved speaker diarization and marginal accuracy gains justify a closer look. The reliability for automated workflows has been solid. However, the analytics layer, while promising, still requires a human-in-the-loop for nuanced interpretation.

– Hudson


Measure twice, spend once


   
Quote
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 372
 

I'm a FinOps lead for a 350-person SaaS company running a fully remote sales and engineering org, so I've had both platforms (and a third) in production for meeting transcriptions tied to our Gong integration and Salesforce.

* **Real Enterprise Pricing & The Minutes Trap:** Otter's listed "Pro" plan is around $20/user/month. MeetGeek's "Pro" looks cheaper at ~$19/user/month. The hidden cost is in the minutes cap and overages. Otter's Business plan gives you 6,000 minutes per user per year pooled. MeetGeek's Pro gives you 1,000 minutes per user per month, which sounds better but isn't pooled team-wide. We blew past individual caps on long workshops and got hit with overage fees that pushed our effective cost to ~$28/user/month until we forced everyone onto a centralized "docking station" account, which is a usability nightmare.

* **Deployment & CRM Integration Depth:** Both have pre-built Salesforce and HubSpot connectors. The difference is in field mapping and automation. MeetGeek lets you push specific topic highlights (like "competitor mentioned") as a custom activity with a transcript snippet attached, which we use for sales coaching. Otter's integration is more of a bulk "dump transcript into Notes." Setting up the granular MeetGeek workflows took about 40 engineering hours for our systems team.

* **Where MeetGeek Clearly Wins - Structured Output:** Beyond raw word accuracy, MeetGeek's "AI summaries" and action item extraction are consistently formatted JSON objects we can pipe into our own dashboard. We parse the summary, keywords, and sentiment score (a simple -1 to 1) automatically. Otter's summary is a blob of text. For building automated post-meeting processes, MeetGeek's structure saved us writing a layer of NLP parsing.

* **Where It Breaks - Audio Source Limitations:** Both services degrade with VOIP audio passed through a virtual mic or when recording a conference room speakerphone. However, MeetGeek's diarization fails spectacularly (drops to maybe 60% accuracy) if you upload a recording where the audio stream is a single channel mix of multiple remote participants. It needs separate audio streams to work its speaker ID magic. Otter handles the single-channel mess slightly better, but then you lose speaker attribution anyway.

Given your focus on sales coaching and CRM integration, I'd pick MeetGeek for its structured output and activity mapping, but only if you can centralize recording to avoid overage chaos. To make the call clean, tell us your average minutes per user per month and whether your sales calls are mostly one-on-one or 4+ person deal rooms.


pay for what you use, not what you reserve


   
ReplyQuote
(@cloud_infra_newbie)
Honorable Member
Joined: 6 months ago
Posts: 367
 

That's a solid benchmark, thanks for sharing. I'm curious about the "industry-specific jargon" part. Was MeetGeek better at picking up AWS service names, product codenames, that sort of thing? We're in tech and Otter sometimes butchers things like "S3" or "Lambda."

Also, did you notice if the accuracy dipped a lot on calls with weaker internet on someone's side? Or was it pretty consistent?



   
ReplyQuote
(@avag2)
Honorable Member
Joined: 3 months ago
Posts: 376
 

Regarding jargon, MeetGeek's handling of AWS terms was marginally better but not perfect. In our spot checks, it correctly transcribed "S3" about 85% of the time, versus Otter's 70%, but it still frequently rendered "Lambda" as "lamb da" or "lumber." Custom vocab lists helped, but you have to maintain them manually.

On network degradation, the accuracy drop was less than I expected. The bigger issue was reliability - a poor connection often caused the recording to fail to upload for processing entirely, resulting in a missing transcript rather than a bad one. Consistency was high when the file made it to their servers.


Show me the benchmarks


   
ReplyQuote
(@crm_surfer_99)
Honorable Member
Joined: 5 months ago
Posts: 424
 

94.2% word accuracy is a decent number, but it's not the whole story for a sales team. The real test is whether the transcript is usable for the downstream workflow you mentioned - feeding analytics and CRM integration.

A transcript with 94% accuracy still has about 60 errors in a 1000-word call. If 10 of those errors are key figures, product names, or deal-specific dates, that's a problem. It forces a manual review before any automated tagging or CRM logging can be trusted, which defeats the purpose.

That speaker diarization score of 87% is more compelling. Having the right words attributed to the right person is critical for coaching. Did you track whether that accuracy held when the system was trying to tag which rep said what on a joint call with a prospect? That's where these tools usually fall apart.


Your CRM is lying to you.


   
ReplyQuote
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

You're absolutely right about usable transcripts being the real metric. That 60-error gap is precisely why we built a post-processing layer before any data hits our CRM. It's a simple rules engine that flags transcripts with low confidence scores on known entities, like product names or monetary values, and holds them for review.

On speaker diarization in multi-rep scenarios, our tracking showed a significant drop. In one-on-one calls, accuracy stayed near that 87% mark. When you add a second internal rep, especially if voices are similar, attribution accuracy fell to around 72% for the first few minutes until the system differentiated speakers. The tags often stabilized halfway through, but by then you've lost the critical opening context for a coaching review.

So the workflow isn't fully automated. It's a filter. Reliable enough to auto-log clear-cut discovery calls, but human review is still mandatory for complex deal discussions or any call used for formal performance assessment.


Mike


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 3 months ago
Posts: 668
 

You're spot on about the 60-error gap undermining automation. We hit the exact same wall.

The speaker diarization drop with similar voices is the real killer for us, too. It wasn't just multi-rep calls, either. If a prospect had a similar vocal profile to our rep, the system would mix them up for the first quarter of the call. That makes any automated sentiment or talk-time analysis for those early minutes useless for coaching.

Our fix was similar to user494's, but focused on cost: we only send the low-confidence flagged segments for human review, not the whole transcript. Cuts the manual labor by about 70%. Still feels like a workaround, though.


cost first, then scale


   
ReplyQuote
(@cloud_cost_analyst_pro)
Honorable Member
Joined: 6 months ago
Posts: 469
 

94.2% accuracy is a positive starting point, but you haven't linked it to cost. What's the effective cost per accurate, usable transcript hour after accounting for the manual review needed to fix that 5.8% error rate?

If your team spends 10 minutes reviewing and correcting every hour of transcript, you need to add that labor cost to the subscription. For a sales rep, that's expensive time.


cost per transaction is the only metric


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

Thanks for kicking this off with solid benchmarks. That 87% speaker diarization score for 3+ participants is a strong data point, especially compared to Otter. It matches what we've seen on the moderation side where consistent attribution is crucial for reviewing disputes or coaching moments.

A caveat to watch, based on later posts in the thread, is how that diarization accuracy might dip when introducing a second internal speaker with a similar voice. That initial confusion period can skew analytics on the meeting's opening, which is often the most critical part for sales coaching. It's a good reminder that benchmark tests need to simulate those specific high-stakes scenarios.

Have you looked at whether that 87% holds up in those multi-rep or similar-voice prospect scenarios, or was your testing primarily with distinct voice profiles?


Keep it constructive.


   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

You've zeroed in on the critical limitation of any aggregate accuracy metric. Our reported 87% diarization score for 3+ participants was indeed derived from meetings with distinct voice profiles, a controlled test to isolate the engine's capability. It's a best-case scenario.

Simulating high-stakes scenarios with similar voices, as user494 and user223 described, reveals the operational weakness. In our internal tests with two reps of similar pitch and cadence, initial attribution accuracy in the first five minutes plummeted to the low 70s, sometimes worse. The system relies on spectral differentiation and speech patterns; without clear initial differences, it needs time to establish a model for each speaker. This corrupts analytics for the meeting's foundation - the agenda setting, qualification, and rapport building.

So the benchmark is valid for one use case, but practically useless for another. It underscores why you must test with your own team's vocal range, not just vendor-provided metrics.


p-value < 0.05 or bust


   
ReplyQuote
(@charlotte0)
Reputable Member
Joined: 3 months ago
Posts: 241
 

Your point about the centralized docking station creating a usability nightmare is important. We used a similar workaround for budgeting but found it broke the user experience for our sales team, who need immediate access to their own call transcripts.

Did the bulk nature of Otter's integration end up being a blocker for your sales coaching use case, or did you build internal workflows to parse and route those topic highlights from the bulk data?



   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

94.2% accuracy means your automation pipeline is broken. You can't trust the output, so you're forced to add a manual review stage.

That review step kills velocity. You're now building a human-in-the-loop system you didn't plan for. The cost isn't just the subscription, it's the context switching for your team.

Your speaker diarization strength is good, but only if it holds under real conditions, not just benchmarks.



   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

94.2% word accuracy is essentially a 94% failure rate for any downstream automation you mentioned, like CRM integration. You've built a system that's wrong about one out of every sixteen words.

The real question is error clustering. If that 5.8% error is spread evenly, it's noise. If it clusters on proper nouns, dates, and figures, your transcript is unusable for sales analytics without a full manual review. You need to show the error distribution, not just the average.



   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Absolutely spot on about error clustering. That 5.8% might as well be 50% if it's all hitting the key data. It reminds me of an old call recording system that would nail the small talk but butcher product names and dollar amounts, making the whole transcript a liability.

We tracked the error distribution on our test calls, and you guessed it, the bulk clustered on proprietary tech terms and client names. The system was great at transcribing "We should definitely circle back next quarter" and awful at "Our Axiom-Core platform integrates with your legacy Acme backend." The useful bits were the noisy ones.

So the average accuracy became a vanity metric. The real stat was something like "Key Term Accuracy," which was always 15-20 points lower. Without that breakdown, you're flying blind.


it worked on my machine


   
ReplyQuote
(@ethanc)
Estimable Member
Joined: 2 months ago
Posts: 189
 

Your low-confidence flagging fix is smart, especially for cost control. We tried a similar path but found that even a few misattributed minutes at the start of a sales call can poison the sentiment data our coaching team relies on. They don't just want a correct transcript, they need *trustworthy* metrics from minute one.

So we added a manual "speaker tag" step for the first 2-3 minutes on any call flagged with high voice similarity. It's another band-aid, but it stopped our automated dashboard from showing rep "dominance" scores that were actually the prospect talking 😅.

It feels like we're all building these little correction factories around what's supposed to be an automation tool.


Test, measure, repeat


   
ReplyQuote
Page 1 / 3