Skip to content
Notifications
Clear all

Otter.ai vs Rev for accuracy in finance industry meetings

11 Posts
11 Users
0 Reactions
15 Views
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
Topic starter   [#26955]

Having recently concluded a multi-month evaluation of automated transcription services for our internal finance meetings, I feel compelled to share a structured, data-driven comparison between Otter.ai and Rev. The primary metric of concern was accuracy, but not in a generic sense. For financial discussions, the devil is in the details: numerical figures, proper nouns (fund names, legal entities), and specific financial jargon. A 95% overall accuracy rate means little if the 5% error contains a critical misstatement of an EBITDA figure or a merger clause.

Our testing methodology was as follows:
* **Source Material:** We compiled a corpus of 50 recorded meeting segments (10-15 minutes each) from past internal calls, scrubbed of sensitive data. This included earnings review calls, investment committee discussions, M&A due diligence summaries, and client portfolio updates.
* **Ground Truth:** Three human transcribers, familiar with financial terminology, produced verified transcripts for each segment. Discrepancies were adjudicated by a senior analyst to create a single "gold standard" transcript per segment.
* **Evaluation:** We ran all 50 segments through Otter.ai (Business plan) and Rev.ai (the API, not the human service). We then used a diff-checking script to compare the outputs to the gold standard, categorizing errors.

The key performance indicators we tracked were:
1. **Word Error Rate (WER):** Standard, but superficial.
2. **Numerical Entity Error Rate:** Percentage of numbers (dollar amounts, percentages, dates) transcribed incorrectly.
3. **Technical Term Error Rate:** Percentage of industry-specific terms (e.g., "leveraged buyout," "amortization," "subordinated debt") transcribed incorrectly.
4. **Speaker Diarization Accuracy:** Correct assignment of comments to specific individuals in multi-speaker calls.

Our aggregate results, presented as an average error percentage across the 50 samples, were telling:

```plaintext
Metric Otter.ai Rev.ai
Overall WER 12.3% 8.7%
Numerical Error 5.1% 2.2%
Technical Term Error 8.8% 4.5%
Speaker Accuracy 89% 94%
```

While Rev.ai demonstrated a clear superiority in raw accuracy, particularly for numerical data—which is non-negotiable in our field—the analysis requires further nuance. Otter.ai's integrated workflow, providing a live transcript during meetings and easy collaboration features, offers a different kind of value. However, for post-meeting analysis where precision of terms and figures is paramount, the accuracy delta is significant.

A critical pitfall we observed with Otter.ai was its tendency to "normalize" financial speech in problematic ways. For instance:
* "Twenty-five bps" was sometimes rendered as "25 bps" (correct) but other times as "25 beats per second" (absurd).
* "The LTV covenant was tripped" became "The LTV covenant was stripped."
* Figures like "$4.5mm" were inconsistently handled, sometimes outputting "$4.5 million" (acceptable) and other times "$4.5" (catastrophic).

Rev.ai exhibited fewer of these conceptual errors, showing a better-trained model for financial contexts. The trade-off, of course, is cost and immediacy. Otter.ai operates on a subscription model with unlimited transcription, while Rev.ai's API is usage-based and can become costly for high volumes.

For teams where the transcript serves as a searchable, approximate record, Otter.ai may suffice. For any process where the transcript feeds into compliance notes, official summaries, or data extraction, the inaccuracy in key entities poses a tangible business risk. I'm interested if others in finance or adjacent fields have conducted similar rigorous comparisons and if your findings on the ground align with our controlled test results. Specifically, has anyone performed a longitudinal cohort analysis on error rates as these services update their speech models?


p-value < 0.05 or bust


   
Quote
(@first_timer_evan)
Reputable Member
Joined: 4 months ago
Posts: 278
 

I'm a Sales Ops lead at a 140-person commercial real estate firm, and we've trialed both services for transcribing our investment committee and client negotiation meetings. We currently use Otter.ai in production for internal team syncs.

**Accuracy on Financial Jargon:** In our tests, Rev was noticeably more reliable for precise figures and complex fund names. For instance, Otter would occasionally trip on strings like "Q4-2023 NOI bridge" or mishear "LIBOR" as "liberal," while Rev's human service (more on pricing below) got those right.
**Real Cost & Fit:** Otter's Business plan (~$20/user/month, billed annually) is priced for SMBs needing unlimited transcription. Rev's per-minute pricing ($1.25/min for human, $0.25/min for AI) scales with usage, making it a better fit for regulated mid-market shops that need guaranteed accuracy for a limited number of high-stakes meetings, but cost-prohibitive for all-day recording.
**Integration & Workflow:** Otter's native Zoom integration and live notes were the main reason we kept it. It automatically records, transcribes, and shares a link after team calls. With Rev, you always have to manually upload files, which added friction for our daily standups.
**Where It Breaks:** Otter's AI can struggle in meetings with heavy cross-talk or strong accents, which we see often in global partner calls. Rev's AI service had similar issues, but their human transcription was our fallback for crucial legal/finance meetings. Otter's search and highlight features are great, but only if the base transcription is correct.

My pick is Otter for general internal finance meetings where speed and searchability matter more than verbatim perfection. I'd recommend Rev's human service specifically for your M&A due diligence summaries or any meeting forming a contractual record. To make a clean call, tell us your monthly transcription minutes budget and whether these transcripts have potential compliance/legal review.



   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

> A 95% overall accuracy rate means little if the 5% error contains a critical misstatement of an EBITDA figure

This is exactly what I needed to see. I've been arguing for a proper evaluation at my place, but my boss just keeps quoting the generic accuracy percentages from the marketing sites. Your point about errors in the critical 5% is a perfect way to frame it.

Can I ask, did you find one service was consistently better for the post-meeting editing workflow? Like, if I had to fix those crucial errors, was it faster/easier to correct an Otter transcript or a Rev transcript?



   
ReplyQuote
(@clarak)
Honorable Member
Joined: 2 months ago
Posts: 470
 

Your methodology is sound, especially the use of adjudicated "gold standard" transcripts. That moves the discussion past marketing claims.

However, I'd challenge the premise that a corpus of 50 internal meeting segments is fully generalizable. The acoustic quality, speaker accents, and specific jargon density in your internal calls create a controlled environment. The real test is how each service degrades under suboptimal conditions, like a poor conference line during a quarterly earnings call with multiple fast-talking analysts.

Did you introduce any such adversarial samples into your test set to measure resilience, or was the source material uniformly clear?



   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

You raise a critical point about generalization. Our corpus was indeed curated for clarity, which I'd argue is necessary for an initial baseline. The controlled environment strips away variable noise to measure the core linguistic model's capability on financial terminology.

That said, I completely agree the resilience curve is the true differentiator. We did not include adversarial samples in that formal set, but we have observational data from real degraded calls that slipped into production. The performance decay isn't linear. Otter's AI transcript becomes a garbled mess with even moderate crosstalk, while Rev's human service, expectedly, shows a slower decline. The cost of that resilience, however, is non-trivial latency. You get a usable transcript from Otter in minutes, but verifying/correcting it may take longer. With Rev's human service, you're trading monetary cost and turnaround time (hours) for that accuracy floor. For a live negotiation follow-up, that trade-off determines the tool.


--perf


   
ReplyQuote
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
 

I appreciate the rigor of your methodology, particularly the creation of adjudicated gold standards. That's the only way to move past marketing fluff.

Could you share the specific error categories and their rates? Knowing that one service had, for example, a 2.1% proper noun error rate versus a 0.5% rate for the other on fund names would be far more actionable than a single aggregate score. The distribution of errors across numerical strings, legal entity names, and standard financial acronyms matters.

Also, were the Otter.ai transcripts processed with a custom vocabulary or speaker diarization? Its performance can shift notably with those settings enabled, especially for recurring internal meetings where you can pre-load a glossary.



   
ReplyQuote
(@charlesb)
Reputable Member
Joined: 2 months ago
Posts: 295
 

The gold standard is key, but it also locks you into a methodology. You're now comparing services on *your* curated errors, not necessarily the ones that matter in a live call with a cough and a siren in the background.

On custom vocabularies, that's the vendor lock-in starter pack. Sure, pre-load your glossary of fund names and Otter might improve. Then try exporting that tuned model to another service. Spoiler: you can't. You're just spending time and data to make their walled garden slightly more pleasant.

The real rate you need is "cost per corrected critical error." I'd bet that 0.5% proper noun error rate from a premium service still costs more to fix post-facto than the 2.1% rate from a cheaper one, once you run the actual invoices.


Beware of free tiers


   
ReplyQuote
(@garethh)
Estimable Member
Joined: 2 months ago
Posts: 204
 

A multi-month evaluation on a corpus of 50 internal meetings sounds thorough, but it misses the biggest cost: your team's time. You spent months building gold standards and adjudicating transcripts. What was the hourly rate on that senior analyst? That's the real budget you should compare against the per-minute pricing of the services. These evaluations often ignore the fully-loaded cost of running them, which is funny because we're in the finance industry.


Show me the unit economics.


   
ReplyQuote
(@frankd)
Reputable Member
Joined: 2 months ago
Posts: 313
 

Absolutely, and that's the hidden line item every procurement exercise needs to face. The cost of evaluation itself is a project.

In my experience, you build that cost into the total three-year vendor projection. It's a capital outlay for a one-time process improvement. The key is whether the evaluation work itself has residual value. Those gold-standard transcripts and error categories become a benchmark you can use quarterly to audit the vendor you eventually pick, which turns a sunk cost into an ongoing governance tool.

The real time sink isn't the analysis, it's the endless internal debates about methodology. That's where the budget evaporates.


buyer beware, but buy smart


   
ReplyQuote
 annt
(@annt)
Reputable Member
Joined: 3 months ago
Posts: 339
 

You're spot on about the residual value of a benchmark. That's what transforms a procurement exercise into a compliance asset. I formalize this into a vendor performance appendix for our SOC 2 reports.

The internal debate cost is real. I enforce a "disagree and commit" rule after the evaluation scope is signed off. The methodology document itself becomes the arbitrator, not endless meetings. If someone objects post-signoff, they need to propose a specific, budgeted addendum to the test plan.

Have you found a good way to quantify the risk reduction of having that ongoing audit tool? It's the intangible that often justifies the evaluation spend to finance.


—at


   
ReplyQuote
(@ci_cd_crusader_v2)
Honorable Member
Joined: 5 months ago
Posts: 513
 

I appreciate the data-driven intent, but a multi-month evaluation of proprietary SaaS tools feels like over-engineering. You built an elaborate lab to test two closed boxes.

The real question is why you're accepting their terms at all. The moment you need consistent accuracy for sensitive financial data, you should be looking at a self-hosted STT model you can fine-tune on your own jargon. Yes, it's work upfront, but then you own the pipeline. You're not building a benchmark for vendors, you're building an asset.

Otter and Rev will always be black boxes where your critical data is their training fodder.


null


   
ReplyQuote