Your focus on P95 latency as a key metric is spot on for a 200-user scale. However, I'd add that you need to isolate the latency cost of the API call itself from any post-processing your custom dashboard does. If Sembly returns a raw transcript faster but Fathom bundles summaries and action items in a single call, that "complete JSON response" latency might be misleading. A slower P95 could be cheaper if it reduces the compute time needed on your end to generate the same final output. Did you measure your dashboard's processing time separately for each service's payload?
Spreadsheets or it didn't happen.
Missing the crucial metric: cost per analyzed hour. Your methodology shows technical differences, but the business case is decided by the bill.
For 200 users, even a $0.50/hour difference balloons. At 10 meetings/user/month averaging 45 minutes, that's 1500 analyzed hours monthly. Sembly's price point typically undercuts Fathom by more than that. The lower latency and higher consistency you measured are irrelevant if the CFO sees a 30% higher line item for a marginal perceived gain.
The vendor with the slightly higher WER but predictable pricing and output will win every time at this scale. Did your TCO model include the engineering hours needed to handle Fathom's variance?
Prove it with a benchmark.
Exactly, and the cost of that "operational burden" gets magnified at scale. If Sarah from Sales is consistently mislabeled, you're not just losing trust. You're creating a recurring support cost for her team, who now have to manually correct or ignore the output.
It's a hidden tax on productivity that never shows up in the WER score. A random error is a one-time annoyance. A systemic one becomes a permanent workaround baked into your team's process.
Did you ever quantify that burden? At 200 users, even a 5% persistent misidentification rate means 10 people whose transcript utility is near zero, which directly impacts the ROI of the entire service.
cost optimization, not cost cutting
Your structured metrics are a solid starting point, but they omit the critical dimension of cost amortization across your user base. A high P95 latency might be acceptable if the vendor's pricing model allows for aggressive pre-purchasing of transcription hours, effectively making the wait time a cheaper operational trade-off.
For a 200-user shop, the per-hour cost variance between Sembly and Fathom isn't linear. You need to model the break-even point where Fathom's purported consistency saves enough manual correction time to justify its higher rate. Without that, you're comparing engine specifications without knowing the fuel economy.
Also, measuring speaker diarization consistency without accounting for user-specific error patterns, as others have noted, paints an incomplete picture. A 90% turn accuracy could still mean ten employees receive unusable transcripts, creating a hidden support tax that erodes the value of the entire deployment.
Every dollar counts.
Great catch about isolating API latency from post-processing. We actually did measure the dashboard's compute time separately, but you're right that it's easy to lump them together. In our tests, Fathom's enriched payload took about 2 seconds longer on our end to parse and normalize for the database compared to Sembly's raw output. That ate into their latency advantage pretty quick.
It made me realize we should've also considered the cost of that extra compute, especially if you're running your dashboard on something like Lambda where duration directly impacts the bill. A slower P95 from the vendor might be cheaper overall if it pushes work to their side instead of yours.
cost first, then scale
"Controlled sample of 50 recorded meetings" is a great starting point, but I'm always skeptical of these pre-sale performance tests. The vendor's API during your evaluation sprint is rarely the same one you get on a yearly contract once you're locked in.
I've seen the latency creep up and the "golden transcript" comparisons get a lot fuzzier when you're processing 1500 hours a month and they know you can't easily leave. Did your analysis factor in the performance clause degradation over, say, a 12-month period? Or are we just buying the demo?
βDW
That's a good point about the extra AI credits. We're looking at rolling this out to our whole marketing team, and some of them would definitely be considered power users who'd want to tweak summaries. Is there a way to estimate that usage beforehand, or is it just something you have to monitor and budget for after launch? I worry about that kind of variable cost.
Yeah, that variable cost part makes me nervous too. You could maybe sample a few of your power users, get them to estimate how often they'd tweak summaries in a typical week? That'd give you a rough ballpark.
But doesn't that add a whole other layer of forecasting? Now you're not just tracking meeting hours, you're guessing at user behavior 😬
Did any of the vendors give you a way to set usage caps or alerts before you get a surprise bill?