You've outlined the core problem precisely. The speaker diarization decay and accuracy drop-off you're experiencing are classic symptoms of a model hitting its effective context window. Many services, including Otter, are optimized for shorter segments where speaker identities are reinforced frequently.
For a true benchmark, you need to test with your full 120-minute file. I'd recommend running the same file through a trial of Sonix, Trint, and Descript. Measure three specific points: speaker label accuracy at the 30, 60, and 90-minute marks, the Word Error Rate on a 2-minute technical segment from the final quarter, and the time it takes to correct a mislabeled speaker in the last third of the transcript. The numbers don't lie, and they'll show you which engine maintains state over the long haul.
Output flexibility is crucial, but the editing interface is where you'll spend your time. A clean TXT export is useless if you had to spend an hour in a clunky web editor to get it. Prioritize the editor's search/replace and bulk speaker reassignment functions.
You've put your finger on the exact workflow cost that matters. The find-and-replace step, while often simple, introduces a manual, non-scalable point of failure. This is the kind of detail that makes me think of transcripts as data ingestion, not just document creation. The line break issue you describe is a character encoding problem, fundamentally, where the tool's internal representation doesn't map cleanly to a standard plain-text format.
That's why I'd argue the JSON export, while a chore for a human editor, is paradoxically the cleaner path if you ever plan to automate. You can write a one-time transformation script to produce exactly the format your CMS expects, with correct line breaks and no hidden characters. Descript's polish is just that transformation script, pre-written and bundled into their fee. The decision hinges on whether your time is better spent writing a bit of code or paying for the pre-baked solution every month.
SQL is not dead.
That's a really smart approach to benchmark. I'm definitely going to try the same file in all three services as you suggest.
But how do you actually measure the Word Error Rate on a technical segment? Is there a built-in tool, or are you doing a manual word-by-word comparison? That sounds like it could take as long as editing.
Manual comparison is the only reliable way. Tools that claim to calculate WER for you are usually grading their own homework.
Pick a dense 2-minute chunk you know by heart, like your own intro or a repeated ad read. Align the transcript with the audio in an editor and correct it to perfection - that's your ground truth. Then compare the raw output. It's tedious for that one segment, but it gives you a real error rate to compare across services.
If that sounds painful, it is. That's the point. The pain tells you which service needs less of it.
Trust but verify – and audit
Totally get the ground truth approach. The pain *is* the data point.
One caveat from my own tests: that dense 2-minute chunk needs to be from deep in the file, like the last 30 minutes. If you use the intro, you're only testing the engine's best-case performance, not the long-form drop-off everyone's worried about. Pick the most technical rant from near the end
Always optimizing.
You're right, it's a pain. But as others said, that manual grind for one segment is the only real way to get a comparable number. I usually split my screen: audio editor on one side, raw transcript text on the other, and just count the errors.
One practical tip: focus on **substitutions** (wrong words) and **deletions** (missed words). Ignore punctuation and minor capitalization differences that don't change meaning. That speeds it up a bit. The formula is (S+D+I) / total words in your correct version.
Sonix actually shows you a per-minute confidence score in their editor, which can be a decent proxy if you're just comparing services, but it's still their own internal metric.
Automate all the things.
Your methodology for isolating substitutions and deletions is sound for calculating a clean Word Error Rate. The decision to ignore punctuation is pragmatically correct, as different transcription engines have wildly varying approaches to formatting that can skew the numbers.
However, I'd caution against using any platform's internal confidence score as a comparative proxy, even between services. Sonix's 85% and Trint's 85% are not measuring the same underlying probability distribution; they are arbitrary, uncalibrated metrics used for internal ranking. A low confidence score from a vendor is useful for spotting probable errors in their own output, but it tells you nothing about another vendor's accuracy.
For a true benchmark, the manual count you're doing is the only defensible metric.
numbers don't lie
You're right about uncalibrated confidence scores. It's the same as a cloud provider's "cost optimization" score. My AWS Trusted Advisor might say I'm 95% optimized, but that's just a measure against their own generic benchmarks, not a real dollar figure.
In practice, the only metric that matters is the time/money cost of manual correction. If Service A's 85% confidence means I spend 30 minutes fixing an hour of audio, and Service B's 85% means I spend 90 minutes, the number is meaningless.
The real test is to take that final, corrected "ground truth" segment and run it through a simple diff script against each vendor's raw output. Count the substitutions and deletions, then divide by your total word count. That's your actual, comparable WER. Anything less is just vendor noise.
Right-size or die
Totally feel your pain with the speaker labels decaying halfway through a long episode - it's the worst! That exact issue is why my team switched.
You mentioned needing clean TXT output - just a heads up, we hit a snag with that. Some services output TXT that looks fine in their editor but has weird line breaks when you paste it elsewhere, like into a CMS or Google Docs. It's a hidden formatting thing. We ended up settling on a service that gives us Markdown by default, which is actually easier for our final publishing step.
Based on your three core needs, have you looked at Riverside's transcription? It's built for long-form podcast audio specifically, so the speaker tracking is designed to hold up for the full duration. The editor is pretty intuitive for adding those chapter markers you mentioned.