Your question about comparing reliability gets to the core of the operational risk. From a stability perspective for automated pipelines, we haven't seen this as an industry-wide hiccup. We use Otter in a parallel workflow for redundancy, and its accuracy, while not perfect, has remained statistically consistent over the last quarter. The key difference isn't just error rate, but error type and predictability.
Fireflies.ai's recent errors appear more systemic - like the "Azure" to "Asana" substitution mentioned earlier - which suggests a model change affecting proper noun recognition. That's far more damaging to automated processes than the occasional filler word mistake. A local Whisper setup provides consistency, but introduces its own pipeline maintenance overhead and lacks the integrated search features that make these SaaS tools valuable in the first place.
So it's not a blanket transcription problem. It's a specific degradation in a service that teams may have integrated deeply, believing its performance was a stable benchmark. The trap isn't just in data lockout, but in architectural lock-in based on a performance profile that no longer exists.