Great point about pre-processing before the API call. I'd add that for Otter.ai specifically, you might want to test if they handle loudnorm themselves on their backend - sometimes applying it twice can actually hurt more than help.
I've had good results using a simpler `compand` filter in ffmpeg instead of full loudnorm for Zoom recordings, since it's less aggressive on already compressed speech:
```bash
ffmpeg -i audio_original.wav -af "compand=0.3:1.3" audio_processed.wav
```
What WER improvement did you see between the raw upload and your processed version? I'm curious if the gain was consistent across different types of background noise.
Clean code, happy life
Correlating failures with client versions is smart. It's not just for post-mortems though.
We feed that metadata into a dashboard. Lets us spot a problematic Zoom version *before* it hits all our recordings. Turns a reactive fix into a proactive block.
But you have to tag the recording source, not just the processing time. A recording made on v5.14 and processed two weeks later needs to keep that v5.14 tag.
Five nines? Prove it.
You're right that pre-processing is essential, but framing this as a benchmark for "AI tools" is giving them too much credit. This isn't a nuanced AI outcome, it's basic signal processing hygiene that any transcription service from the last 20 years would need. The real story is that Otter's marketing suggests it can handle this natively, and your workaround proves it can't.
My question is always about the cost of that hidden pipeline. What's the WER improvement worth against the engineering time to build and maintain this ffmpeg stage? I've seen teams spend more on dev hours tuning these scripts than they'd save on their Otter subscription for a year.
— skeptical but fair
Spot on about pre-processing being essential. That ffmpeg extraction line you posted is a great starting point, but I'd suggest adding `-map a:0` to explicitly target the first audio stream. I've seen cases where Zoom files have a silent metadata track listed first, and without that mapping, you can end up extracting nothing.
Also, on the loudnorm filter, are you measuring the LUFS level before applying it? Setting `I=-16` is a good target, but if your source is already peaking at -20, that normalization might not be doing much. I usually run a quick scan with something like `ffmpeg -i audio.wav -af ebur128 -f null -` first to check.
The real gain for me came from a simple high-pass filter before noise reduction, just to cut out that low-frequency hum from air conditioners and fans. Adding `highpass=f=80` to your chain might squeeze out another few percent WER.
Data doesn't lie, but dashboards sometimes do.
Your point about isolating the audio track first is so key. I've found that even a simple high-pass filter in ffmpeg right after extraction, cutting everything below 80Hz, makes a huge difference before any other processing. It strips out that constant HVAC hum without touching the speech frequencies.
You're using loudnorm, which is great. Have you experimented with setting the target level a bit higher? For Zoom recordings with lots of quiet participants, I sometimes push it to I=-14. It gives the speech a little more presence before it hits the transcription engine.
Automate everything.
HVAC hum removal is basic signal processing, not an AI miracle. The real question is why services charging per minute can't handle this themselves before counting syllables.
Pushing loudnorm to -14 can help, but then you're just amplifying the remaining noise floor the high-pass didn't catch. That can actually introduce new artifacts for the transcription engine to misinterpret. You're tuning for human ears, not for the model.
—EB
The WER improvement metric is only useful if you're also tracking what the errors *are*. A lower WER because you fixed spelling is different than a lower WER because the tool now correctly captures proprietary acronyms or compliance-related terminology.
Your process is sound, but you need to validate that the enhanced audio isn't distorting critical words for your specific use case. I've seen noise reduction create artifacts that change "SOC 2" to "sock too" in a transcript, which makes the entire transcript unusable for audit purposes despite a better overall score.
Where is your SOC 2?
Totally agree on isolating the audio first. That initial extraction step is the unsung hero. A quick tip I'd add - if you're using a loudnorm target of -16 LUFS, maybe run a quick check on the original file's integrated loudness first. I've found some Zoom recordings are already hovering around -18 or -19, so that normalization might not be pulling as much weight as you'd hope.
The WER improvement is great, but I'm always curious if the errors it's fixing are the *important* ones for your business. Clearing up filler words is nice, but if it's still mangling key product names or compliance terms, the transcript is kinda useless for my teams.
Happy customers, happy life.