Voice fingerprinting fails with scale? Shocking.
> the only reliable method is to manually edit the transcript
You've discovered the actual feature. The pitch is automation. The product is a glorified text editor with timestamps. Your 20-minute cleanup chore *is* the workflow.
What if you're wrong about needing to track every action item to a person? For large groups, the useful record is often 'what was said,' not 'who said it.' Chasing perfect attribution is where the cost creeps in.
Doubt everything
That's the real shift in thinking most teams miss. If the attribution is truly crucial, then the meeting structure itself is broken - you can't outsource accountability to a bot. The record keeper should be a designated human role, not a feature.
But you're right, the 80% use case for big meetings is capturing the content, not the cast list. The problem is these tools are built and sold on the latter, so you're paying for the wrong thing and still doing the chore. I've seen more teams just use the raw, unattributed transcript from something like Whisper, then paste it into a doc for collaborative notes. That's often faster than "fixing" the output from a tool that promised more.
It's a different approach, but Otter inherits whatever garbage display names people are using. It just swaps "Speaker 3" for "Tim's iPhone" and now you're cleaning that up instead. The identification source changes, not the problem.
Your point about similar vocal ranges is the whole scam. They demo with three distinct voices, but real meetings sound like a homogenous audio soup after the first ten minutes.