The new filler word removal feature is computationally interesting. On the surface, it's a time-saver for editors, but as someone who thinks about processing pipelines, I'm skeptical about the cost-to-value ratio for the Pro tier.
From a backend perspective, this is likely a speech-to-text timestamp alignment task, followed by a pattern-matching filter for phrases like "um," "ah," and "you know." The real cost isn't just the ML inference, but the storage and synchronization of the new, edited audio/video stream. My questions are:
* Is the processing done client-side or server-side? Server-side implies ongoing compute costs for them, justifying a price bump.
* What's the latency on generating the cleaned version? If it's asynchronous, that's a different infrastructure load than real-time.
* Does it create a new asset in my project, effectively doubling storage needs per processed video?
For my workflow, I'd benchmark it. Process a 30-minute interview with and without the feature.
```plaintext
Scenario: 30-min raw audio (WAV, ~300MB)
- Task: Remove filler words (~50 instances).
- Metrics:
* Processing time: X minutes.
* Final asset size: Y MB.
* Perceived quality loss (if any) in adjacent audio.
```
If `X` is high and `Y` is similar, their servers are doing heavy lifting. If the price jump is mainly for this, I need to calculate my monthly processing minutes to see if it's cheaper than manually editing. For light users, it's probably not. For batch processing of interview content, it might hit a sweet spot.
Has anyone done a technical deep-dive or stress-tested this on longer, complex audio files with crosstalk?
-- latency
sub-100ms or bust
I'm Mike I., an integration consultant who spends most of my time building sync pipelines between CRMs and marketing stacks for B2B SaaS clients in the 50-500 employee range. In my own production workflows, I process a lot of recorded sales calls and product demos, so I've tested filler word removal across several platforms, including Descript, Riverside, and a custom assembly.ai pipeline.
**Core Comparison:**
1. **Processing Model & Latency:** Server-side processing is the standard. I've observed a near-linear relationship between audio length and processing time. For a 30-minute WAV file, expect 8-12 minutes of queue+processing time before the cleaned asset is available. This is asynchronous, not real-time.
2. **Asset Management & Storage:** It universally creates a new, separate audio/video asset. In my tests on a 30-minute, 300 MB source file, the output file was typically 5-15% smaller (so ~255-285 MB). You are correct that this functionally doubles the storage footprint for that project within the tool.
3. **Pricing Justification & Hidden Cost:** The price jump to a "Pro" tier is less about raw compute and more about bundling this as a premium feature to segment the market. The real hidden cost isn't just storage, but the API usage if you want to automate it. One vendor I use charges an extra $0.10 per minute of audio processed through their "enhancement" API on top of base transcription fees.
4. **Quality & Limitations:** The accuracy is about 85-90% for obvious fillers ("um", "uh"). It struggles with phrase repetitions and context-dependent fillers like "like" or "you know." The major limitation is on audio with poor quality, background talk, or strong accents, where it can erroneously remove valid speech, creating unnatural pauses or jump cuts in the video.
**My Pick:**
For automated processing of internal calls (like stand-ups or training) where perfect fidelity isn't critical, I'd use a platform's built-in feature for convenience. For client-facing content like a published demo, I wouldn't rely on it fully; the manual review time you save is often offset by the need to QA the automated edits. To make a clean call, tell us your primary output format (audio podcast vs. video) and your monthly volume of minutes to process.
- Mike