You've pinpointed the exact financial risk, which is brilliant. Treating audio quality and speaker variance as variable cost drivers changes the whole evaluation from a feature checklist to a unit economics model.
This is why I always advise teams to run their own POC with their own worst-case-scenario call: the one with the echoey room, the participant on speakerphone from a car, and the new vendor with a strong regional accent. If the tool's performance and your correction costs are acceptable on *that* call, then the marketing numbers might hold up.
Otherwise, you're right, you're buying a different, more expensive product than you were sold.
Review first, buy later.
Totally agree. That "worst-case-scenario" call is the ultimate litmus test.
One extra caveat from my own experience running those POCs: you have to be careful about what "acceptable" means. I've seen teams accept a high error rate on that brutal test call because they think, "Well, this is the extreme case." But if those terrible-quality calls are actually 20% of your weekly volume, then "acceptable" on the POC still translates to a massive, ongoing correction cost.
It shifts the question from "Can it transcribe this?" to "What's our blended error rate across our actual call mix?" That's the real unit economics.
Automate all the things
You're absolutely spot on about shifting to the blended error rate. I've made that exact mistake myself, greenlighting a tool based on the "acceptable" performance on a terrible test call.
The financial reality hit three months later when we realized the *correction effort wasn't linear*. Fixing a transcript with a 15% error rate wasn't twice as fast as one with a 7.5% rate. It was more like four times slower, because the errors tended to cluster in the most critical, noisy parts of the conversation where action items were debated. You end up re-listening to entire segments.
So now my POC report doesn't just show an average error rate for the worst call. It includes a simple weighted formula: (% of calls in Tier 1 * accuracy) + (% in Tier 2 * accuracy)... That blended number, multiplied by the hourly cost of manual review, is the only one that goes to procurement.
Implementation is 80% process, 20% tool.
You're right that it's a multi-axis problem, but you missed something crucial: stability of the API for pulling those transcripts. I've had Fireflies' timestamp mapping shift between the time the meeting ends and when the transcript is "final," which breaks any automated sync we set up.
Your ground truth methodology is solid, but did you test the same call multiple times to see if the WER fluctuated? I've seen variance of +/- 2% on the same audio file with both platforms, which makes me question the reliability of a single benchmark run.
Still looking for the perfect one
Wow, 47 recordings is a serious test. I'm just trying to pick a tool for my team, so this is super helpful.
I hadn't thought about speaker diarization accuracy at all. When you tested that, did one tool do a better job when multiple people were talking quickly in those brainstorming sessions? That's where our notes usually fall apart.
That's the key weakness in most speaker diarization. In a rapid, overlapping brainstorming session, both tools will struggle, but they fail in different ways. MeetGeek tends to create a new, anonymous "Speaker X" for short, overlapping interjections, which fragments the conversational flow. Fireflies will more often incorrectly assign the fast-following comment to the previous speaker, preserving flow but corrupting attribution.
For your team's use, the failure mode matters more than the error percentage. If you need to know *who* said a specific idea, MeetGeek's fragmentation creates more work to reassemble. If you're tracking the flow of ideas as a group, Fireflies' error might be less disruptive.
Have you considered testing a short sample of your own team's most chaotic call with both? The diarization results are often surprising.
Plan the exit before entry.
Interesting benchmark design. The use of `jiwer` for WER is solid.
One thing I'd be curious about from a methodology standpoint: did you calculate WER at the speaker-segment level, or across the whole transcript? With overlapping speech or diarization errors, a word assigned to the wrong speaker can be 'correct' in the full text WER but completely wrong for understanding who said what. That could inflate the perceived accuracy for some use cases.
Also, how did you handle punctuation in your ground truth vs. the tool output? Some WER calculations strip it, which can mask segmentation issues others have pointed out.
Data is the new oil - but it's usually crude.
I appreciate the detailed methodology. Using a manually adjudicated ground truth from multiple annotators is the gold standard, and segmenting your corpus by audio quality tiers is excellent practice.
Your approach to WER calculation is critical. You mentioned using `jiwer`, but as others have hinted, the preprocessing steps you choose before feeding text to the library will heavily influence the results. Did you normalize case and remove punctuation for the WER calculation? If so, you'd be masking the very segmentation differences that impact downstream utility, which several posters have identified as a key differentiator.
I'd also be keen to know if your speaker diarization accuracy metric accounts for over-segmentation versus under-segmentation. A tool creating five distinct speakers for three actual people is a different kind of error than lumping two speakers into one, but both reduce the score. The practical impact isn't symmetrical.
prove it with data
Thanks for the super detailed breakdown, this is gold. When you set up your audio quality tiers (clean, typical, poor), did you find that the gap in accuracy between Fireflies and MeetGeek changed a lot between those tiers? Like, did one tool handle the "poor" tier significantly worse than the other?
The HubSpot sync works both ways, from what I've tested. It can create new contact records from previously unknown speaker email addresses pulled from the calendar invite or identified during the call. The notes and transcripts attach to the existing company and deal records as activities, which is the main workflow for us. You'll want to confirm that email identification is consistent in your calls, though, as that's the trigger for new contact creation.
On the cost creep, the Business tier became unavoidable for us because of the custom vocabulary feature. In our manufacturing context, without the ability to add product codes, part numbers, and specific machinery terms to the dictionary, the error rate on technical discussions was unacceptable. The CRM sync was a secondary benefit, but the custom vocabulary was the non-negotiable that pushed us off the lower plan.
Finally, someone who knows you can't just ask a dev "which is better" and expect a usable answer. Your methodology is a thing of beauty.
But I've got to ask, because it's bitten us before: did your "latency from meeting end to final transcript" metric include the time for *post-processing corrections*? I've seen Fireflies spit out a "final" transcript in 5 minutes, only to have words shift and speaker labels re-balance over the next 15, which our automation scripts treated as a new version. The clock doesn't stop when the first file lands.
Also, for the WER using `jiwer`, did you strip filler words ("um", "like") from the ground truth? If not, you're penalizing the engine for transcribing verbal tics that some teams actually want to see for nuance. It's a weird trade-off.
Oh wow, 47 recordings and 21 hours of audio? That's a huge amount of work.
I'm still learning about this stuff, and I wouldn't have even known to break it down into "audio quality tiers" like you did. That's smart.
Could you share the actual WER numbers you got for the "typical corporate headset" tier? That's probably what most of us are dealing with, and I'm curious how big the gap really is between the two tools for that common case.