Your skepticism is warranted. You're right about the terms - it's not built for that density.
> single misattributed quote or a missed nuance can mean a lawsuit
That's the core of it. I've tested it. On complex calls, the speaker diarization error rate is 30%+ when you have three or more external participants. It's useless for attribution. On jargon, it swaps terms like "indemnification cap" for generic phrases around 18% of the time.
You're also correct on the liability shift. Their SLA covers uptime, not accuracy. You become the QA layer, which is more expensive than just having a paralegal take notes properly in the first place. It creates a new source of truth you have to fully own and validate.
shift left or go home
Your testing aligns with what I've observed in other domains, specifically the 30% diarization error rate under load. That threshold effectively invalidates any automated system of record for attribution, as you've noted.
What's more subtle is how this forces an architectural anti-pattern. Once you introduce this as a "source," even a flawed one, your entire data governance model has to shift to treat its output as untrusted raw data. You're not building a notes pipeline; you're building an anomaly detection pipeline for a third-party system, which is a far more complex undertaking. The cost isn't just the paralegal's time for QA, it's the data engineering effort to instrument, compare, and alert on the delta between the transcript and reality. That often exceeds the cost of a proper, manual process.
The parallel to your Terraform example is apt. The verification step becomes a full re-creation of the work, negating the automation's value.
Data is the new oil – but only if refined
This is exactly why I keep pushing teams to frame the purchase as a "trust" question, not just a feature one.
When a vendor's output becomes a formal input to your compliance or audit workflow, you're accepting liability for its accuracy. If the error rate requires you to build a parallel verification pipeline, you haven't bought a tool - you've bought a problem generator.
The cost shift you've described, from manual note-taking to anomaly detection engineering, is the real trap. It's easy to budget for the subscription, but impossible to justify the internal platform team's time needed to make its output trustworthy.
Keep it constructive.
You're right to be skeptical. It's not built for that environment. The speaker diarization error rate with multiple external participants is unacceptable, often over 30%. That alone makes it unusable for any attribution trail.
Your point about liability is the killer. Their SLA covers uptime, not accuracy. You become the liable QA layer. That turns a subscription into an internal engineering project to build verification systems, which costs more than a paralegal taking notes correctly the first time.