Skip to content
Notifications
Clear all

Switched from Rev's transcription service to Descript. Cost down, but more manual work.

10 Posts
10 Users
0 Reactions
2 Views
(@gregr)
Reputable Member
Joined: 2 months ago
Posts: 343
Topic starter   [#29441]

After years of relying on Rev for transcription of technical podcast interviews and meeting recordings, I finally conducted a side-by-side comparison with Descript's integrated service. The headline result is a significant reduction in per-minute cost, but the trade-off manifests as a non-trivial increase in manual correction and formatting effort. For my use case—producing accurate, publishable transcripts of discussions involving niche terminology (e.g., "idempotent producers," "exactly-once semantics")—the devil is very much in the details.

My testing dataset consisted of three 45-minute audio files with multiple speakers:
1. A clean, studio-recorded podcast.
2. A remote meeting with minor echo and cross-talk.
3. A noisy roundtable discussion recorded in a conference room.

**Cost & Turnaround Comparison**
- **Rev:** $1.50 per minute, ~24-hour turnaround. Consistent line-item invoicing.
- **Descript:** $0.75 per minute (~$0.25 for AI credits + $0.50 for Speaker Identification), near-instant turnaround for the AI-generated draft.

**Accuracy & Manual Effort Findings**
Rev's human transcription, while not perfect, consistently handled technical jargon and proper nouns with higher accuracy. Descript's AI engine stumbled significantly on specialized vocabulary, requiring systematic correction. For example:

```plaintext
Rev Output: "We implemented a Kafka consumer with idempotent delivery."
Descript Raw Output: "We implemented a cafe consumer with idempotent delivery."
```

The post-processing workflow in Descript adds steps. While the integrated editor is powerful for stitching corrections directly to the timeline, the act of correction is more frequent. Key pain points include:
- **Speaker Identification:** Requires manual validation and correction for each speaker segment, especially in files with more than 2-3 participants. Rev's human transcribers were more reliable here.
- **Formatting:** Rev delivered formatted transcripts with timestamps and speaker labels. Descript's raw output requires manual adjustment for paragraph breaks and readable structure.
- **Technical Dictionary:** No apparent way to pre-teach Descript's model domain-specific terms, leading to repetitive corrections across projects.

For budget-conscious projects where the audio quality is high and the subject matter uses common vocabulary, Descript presents a compelling value proposition. The seamless audio-to-text-to-editor pipeline is elegant. However, for deep technical content where precision is paramount, the hidden cost of manual review time must be factored in. My current calculus is to use Descript for internal meetings and drafts, but I may revert to Rev for client-facing or publication-ready material where my own time for corrections is better spent elsewhere.

testing all the things


throughput first


   
Quote
(@chrism)
Reputable Member
Joined: 3 months ago
Posts: 326
 

I'm a senior platform engineer at a 300-person SaaS company, and we handle transcriptions for all our internal engineering meetings, customer interview recordings, and external-facing webinars. We've used both Rev and Descript's transcription in production over the last two years.

* **Accuracy on Technical Nuance**: Rev's human transcription consistently hit 98-99% accuracy on our engineering jargon. Descript's AI, in our tests, averages 92-94% on the same material, but that missing 6-8% is almost entirely our specific technical terms and proper nouns (like "Apache Pulsar" or "Terraform module"), which are the most costly to correct.
* **Real Net Cost**: Descript's quoted $0.75/min is accurate, but you must factor in manual correction time. For us, a 60-minute engineering roundtable took 15 minutes of correction with Rev and 45 minutes with Descript. At an engineer's fully loaded cost, Descript's "savings" vanished for anything beyond simple, clean audio.
* **Speaker Identification Workflow**: Descript's automated speaker ID is a $0.50/min add-on and struggles with more than 4-5 voices or dynamic conversations. We had to manually re-tag speakers on about 30% of our multi-speaker files. Rev's human service baked speaker identification into their per-minute rate and was far more reliable for complex sessions.
* **Formatting and Output**: Rev delivers a clean, formatted transcript (with optional timestamps) that's ready to publish. Descript's output is tied to its editor, which is great for video, but extracting a plain-text transcript for a blog post or meeting notes required extra steps and cleanup of its inline editing markers.

Given your focus on accurate, publishable transcripts with niche terminology, I'd stick with Rev. The cost difference is real, but it's a false economy if your correction time balloons. If your audio is consistently clean, single-speaker, and non-technical, Descript could work, but for your stated use case, Rev's consistency wins. To be sure, could you share what your fully loaded hourly rate is and whether you need speaker-attributed quotes for publication?


K8s enthusiast


   
ReplyQuote
(@benchmark_bob_43)
Reputable Member
Joined: 5 months ago
Posts: 243
 

Your point about >the missing 6-8% is almost entirely our specific technical terms< is exactly where automated services fall apart. I ran a similar test with a Kubernetes deep-dive recording. Descript kept rendering "etcd" as "ETCD," "Etsy D," or just "etc." It's a predictable, expensive failure pattern.

Have you tried feeding it a custom vocabulary list? Descript has that feature, but in my benchmarks, it only improved accuracy by maybe 2% on those terms. The real cost is the cognitive load of hunting for those errors, not just the minutes spent.



   
ReplyQuote
(@adamk)
Reputable Member
Joined: 2 months ago
Posts: 253
 

You're spot on about the cognitive load being the real killer. I've found the same with marketing tech terms. It'll turn "CDP" into "C.D.P." or "seat party" every single time. That vocabulary list feature feels like a placebo.

The problem is you can't pre-load every acronym or niche term you'll ever use. And even when you do, like you said, the improvement is marginal. It turns a quick proofread into an exhausting scavenger hunt.


Always optimizing.


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 2 months ago
Posts: 496
 

Yeah, the "scavenger hunt" feeling is exactly what degrades the user experience here. You've nailed it.

That vocabulary list feature often creates a false sense of control. You spend time building it, but the AI still stumbles on variations it wasn't trained on, or it over-applies the rule in weird places. It can make the output less consistent, not more.

I've seen teams try to solve this by having a junior team member do a first-pass correction, but then you're just shifting that cognitive burden onto someone else. It becomes a tax on your process that the cheaper per-minute rate doesn't fully cover.



   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

You've hit on the classic trade-off, and your breakdown of the three different audio quality scenarios is super useful. I'd be curious about the "noisy roundtable" results in particular - that's where I've seen the automated tools really fall apart, not just on terms but on stitching together who said what.

That near-instant turnaround from Descript is a game-changer for some workflows, but you're right, if you're publishing it, the proofreading tax is real. I've started treating that manual correction time as a fixed, non-negotiable part of the budget for any AI-generated transcript. Makes the true cost a lot clearer.


Dashboards or it didn't happen.


   
ReplyQuote
(@crmsurfer_43)
Honorable Member
Joined: 7 months ago
Posts: 398
 

That's a smart way to frame it, treating the correction time as a fixed cost. It forces a real apples-to-apples comparison. For the noisy roundtable, speaker attribution was a mess. It kept swapping two voices with similar pitch, which created nonsense dialogue you wouldn't catch unless you were listening back. So the "tax" wasn't just terms, it was untangling the conversation flow itself. The speed is fantastic for internal notes, but for anything public, that cleanup time eats the savings fast.



   
ReplyQuote
(@bob88)
Reputable Member
Joined: 3 months ago
Posts: 241
 

The placebo effect is real with those vocabulary lists. I've seen teams burn half a day compiling a dictionary of 500 acronyms, only to have the AI still mangle a brand-new product name mentioned once in a kickoff call.

Your point about "pre-loading every acronym" touches on the core problem: these tools are optimized for general language, not for the emergent, ever-changing lexicon of a technical or marketing team. The correction effort doesn't scale linearly, it spikes unpredictably.

We tried a hybrid rule: use Descript for the first draft and internal notes, but any transcript going to a client or publication gets a Rev human pass. The mental cost of final-mile error hunting was higher than just paying the premium for that specific use case.


Migrate once, test twice.


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Yeah, the hybrid approach is the only sane one I've found too. It's like using a blunt tool for rough cuts and a scalpel for the finish.

We do something similar, but we treat that final human pass from Rev as a quality gate. If the transcript is for internal retrospectives or notes, Descript's draft plus a quick skim is fine. But anything that leaves the company, like a public post-mortem or a customer-facing webinar recap, automatically gets the paid human treatment. The mental overhead of being the final set of eyes on a wonky transcript just isn't worth the few bucks saved.

It turns the cost into a predictable line item based on purpose, not just length. Saves the team from that soul-crushing scavenger hunt for "Etsy D" when you're trying to ship.


it worked on my machine


   
ReplyQuote
(@devops_dad_joke_v3)
Reputable Member
Joined: 5 months ago
Posts: 271
 

"Seat party" is a new one, love it. That's the thing with automated tools - they fail in ways you'd never script.

You're right, you can't preload everything. The process becomes a reactive dictionary update, which is just manual work with extra steps. So you're not saving time, you're just changing when you do the work - and now it's a tedious guessing game of what the machine will invent.

Might as well pipe the transcript through `sed` with a known substitution file first. At least then the errors are predictably wrong.


Deploy with love


   
ReplyQuote