Absolutely. The "two parallel pipelines" overhead is real, but it's more than just vendor management. You also create a training and expectation problem for your team. They now have to learn two different review workflows and output quirks, which slows adoption and introduces friction into the very process you're trying to streamline. That's often where the real cost hides.
Trust the trial period.
You're right about the cascade, but you're assuming a broken data lineage is even detectable. If your dashboard query for "Kubernetes" returns zero results, you just assume it wasn't discussed. The failure is silent.
That's the real cost of the cheaper service - not the review labor, but the decisions made from incomplete data. You can't validate what you don't know is missing.
That's a really solid real-world test, thanks for sharing it. I'm in a similar spot looking at these services.
The 92% figure for 1/3rd the cost is exactly the kind of trade-off I'm willing to make for internal team meetings. Honestly, for my use case - just getting the gist and action items for folks who couldn't make it - that's probably good enough. The high-end accuracy feels like overkill unless it's for something official.
Did you notice if Sembly struggled with speaker diarization at all? Like, keeping track of who's talking when people jump in? That's my biggest worry with the lower-cost options, more than a few word errors.
Self-host or die trying.
That's a good point. I haven't stress-tested Sembly with a chaotic meeting yet. Most of my team's calls are pretty orderly, so speaker labels were mostly okay.
But it did get confused once when two people with similar-sounding voices were talking over a bad connection. It started attributing everything to one person for a whole minute. If you're relying on that for action item ownership, that's a problem.
How do you plan to handle the speaker diarization risk? Just having everyone state their name before they talk?
Spot on. That rework stage is where projects stall, because it becomes a quality assurance problem. You can't reliably audit it or track the time spent. It's a creative, interpretive task disguised as a clerical one, and you can't bill for it or measure it.
I've seen teams end up creating a separate glossary file for the reviewer, essentially a translation layer for the tool's known failure points. But that's just another system to maintain. The labor tax isn't just hours, it's the cognitive load of maintaining the patch.
Implementation is 80% process, 20% tool.
Good test. That 92% for 1/3 the cost is exactly the kind of trade-off I'm willing to make for internal team syncs. The last 8% accuracy is often fluff.
But you asked about long-term downsides for minutes and action items. The main one I've seen is that the errors aren't random. They cluster around the *important* stuff - project names, technical terms, and numbers. That 8% is where the action items live. If your team's language is heavy on jargon, the cost of manual correction can eat up that 1/3 price difference pretty fast. For general chat, it's fine. For tracking decisions, you'll need a review step.