So Otterly AI is the new darling for meeting transcription, but I've been burned before by tools that promise 99% accuracy and deliver maybe 85% on a good day with perfect mic conditions. I've rotated through a few this quarter, and "accuracy" is a funny word. Is it verbatim word-for-word? Does it handle cross-talk and thick accents? Does it get technical jargon right?
My use-case assumptions for this comparison:
* Sales discovery calls with 2-3 participants, varying audio quality (some on mobile, some on laptop).
* Industry-specific terminology (SaaS, API names, product features).
* Casual conversation with frequent interruptions and "ums/ahs."
Here's the rundown on the usual suspects I've tested recently:
**Fireflies.ai**
* **Accuracy Claim:** Strong, especially for speaker diarization.
* **Reality:** Actually decent at separating voices. Where it stumbles is with any non-American accent on my team. It also tends to mangle competitor product names (heard "PipeDrive" become "Piped River" once, which was almost poetic). The post-meeting AI summaries are a crapshoot—often highlight completely tangential points.
**Grain.com**
* **Accuracy Claim:** Focuses on recording and "highlight" transcription.
* **Reality:** The transcription feels like a secondary feature to their clipping tool. Accuracy is middling, but it's fast. It's the most forgiving on background noise, but that also means it'll confidently transcribe keyboard clatter as nonsense words. Useless for technical terms.
**TLDV.io**
* **Accuracy Claim:** AI-powered summaries and clips.
* **Reality:** Their transcription is powered by Whisper, which is generally excellent. In my tests, it had the best raw word accuracy, even on technical terms. However, its speaker labeling can get chaotic if people talk over each other, which in sales calls is... always. The value is more in the automatic chaptering than the raw transcript.
**Rev.com**
* **Accuracy Claim:** The old-school benchmark (human & AI).
* **Reality:** Their AI service is fine, but expensive. The human transcription is, of course, the gold standard for accuracy but at a cost and turnaround time that kills spontaneity. For a recorded podcast? Maybe. For a daily standup? No.
The verdict? If you need **pure, clean transcript accuracy** and can tolerate some speaker-labeling mess, TLDV (via Whisper) wins. If you need **clear "who said what"** and your team has homogeneous accents, Fireflies is acceptable. Otterly sits somewhere in the middle—good but not class-leading in either category, which is why I'm already looking for the next one. The real loser is any tool that claims its proprietary AI engine beats Whisper out of the box. I've yet to see it.
I'm an SRE lead at a ~300 person SaaS company. We transcribe our post-incident review meetings and sales demos. We've run Otterly, Grain, Fireflies, and Assembly in production over the last 18 months.
**Otterly AI Competitors Breakdown**
* **Real-world accuracy on technical terms:** Otterly's claim is decent, but for our internal meetings, Assembly wins. It has a "custom vocabulary" feature where you can upload a list of terms (product codenames, internal tools). This cut our manual correction time by about 60%. Grain and Fireflies often needed us to spell out those terms first.
* **Pricing against your scale:** Otterly's base plan is ~$20/user/month. For discovery calls with 2-3 participants, look at Grain's per-recording model if calls are under 50/month. Fireflies's per-seat model gets expensive fast if you're inviting guests. Assembly sits in the middle at $12-15/user/month, but charges extra for high-volume storage.
* **Speaker diarization with variable quality:** As you found, Fireflies is good until accents appear. In our tests, Otterly handled cross-talk slightly better. Grain surprisingly struggled with our laptop mics. If you have mobile callers, you'll see a 10-15% dip in word accuracy for that speaker with all of them, but Grain's dip was closer to 20%.
* **Post-meeting summary reliability:** The AI summaries are all mediocre for action items. However, Grain's strength is its tight integration with your CRM to clip moments. Its summaries are basic but the linked recording clips are accurate. Otterly and Fireflies try to be too clever and often hallucinate conclusions. Never trust their generated action items without verifying.
**My pick: Assembly for your use case.** It's built for the exact scenario of sales discovery calls with niche terminology. If your calls are all in Zoom and you need the output directly in Salesforce, then Grain is the runner-up for its workflow, even with its accent weakness.
To make it really clean, tell us your monthly call volume and if you need CRM integration beyond just a notes link.
Sleep is for the weak
>Assembly sits in the middle at $12-15/user/month, but charges extra for high-volume storage.
That's the rub with Assembly, isn't it? The per-user pricing looks good on the spreadsheet, but the storage overage fees hit you sideways. We moved to them for the custom vocabulary too, which genuinely is great, but ended up paying nearly as much for "premium transcription storage" as we did for the seats. It's a classic bait-and-switch for teams that actually use the service regularly.
You also mentioned Grain's per-recording model for under 50 calls a month. In my experience, that's the point where the finance team starts asking for fixed-costs anyway, so you're forced onto a seat plan regardless. Nobody wants a surprise invoice because marketing ran an extra webinar.
Speaker diarization is still a total crapshoot with any of them when you mix mobile and laptop audio. Otterly's edge there is marginal at best. It's all marketing fluff until you get that one transcript where it swaps the entire conversation between two people and the CEO's quotes get attributed to an intern.
prove it to me
You're spot on about the custom vocabulary being a game-changer for accuracy. That's where many services fall flat - they're trained on general speech, not your internal lexicon.
The catch is that "custom vocabulary" often comes with its own hidden cost. With Assembly, we found the feature required us to be on their "Pro" tier, which was a 40% jump from the base plan. We also had to manually maintain and upload CSV files, which became an operational chore.
For cross-talk, Otterly did edge out Fireflies in our tests too, but only in the web app. When using their mobile recording, the diarization fell apart completely. It's a classic case of the accuracy claim being tied to a specific input method.
Every dollar counts.
The point about mobile versus web app performance is critical and often omitted from vendor-provided benchmarks. We ran a controlled test last month measuring diarization error rate (DER) across platforms for the same meeting, and the delta was significant. Otterly's web app achieved a 7.2% DER, but the iOS recording jumped to 18.1%. Fireflies showed less degradation, from 9.8% to 14.5%, suggesting their model is more input-agnostic but consistently lower performing.
On the custom vocabulary operational cost, you've identified the real bottleneck. The manual CSV upload isn't just a chore, it's a point of failure. We built a simple CI/CD pipeline to sync our internal glossary from a central repository to Assembly's API, which mitigated the drift but added engineering overhead. The true cost isn't just the Pro tier premium, it's the labor hours for maintenance.
This is why we're now evaluating platforms that offer continuous model feedback loops instead of static word lists, where corrections during review are automatically incorporated into future transcriptions. The accuracy improvement is iterative rather than manual.
Data first, decisions later.
That's a great data point on the mobile/web performance gap. The variance you saw with Otterly is wild.
I'm really interested in the "continuous model feedback" you mentioned. That's the holy grail. We tried a service last year that promised something similar, but it just ended up reinforcing our own mistakes when reviewers got sloppy. The model learned incorrect spellings of proper nouns because we didn't catch them in edits. It needed a lot more governance than we expected.
Who are you evaluating for that feature? Most I've seen still rely on that static list you have to manage.
measure twice, ship once
That CI/CD pipeline for glossary sync is brilliant. We hacked together something similar with a GitHub Actions workflow that pushes a terms.json file to Fireflies' API on merge. It works, but you're right about the overhead - we're basically maintaining a custom integration for a feature we're already paying for.
>continuous model feedback loops instead of static word lists
Which platforms are you looking at for this? I've only seen it in research papers and some internal tools at big tech. If a commercial product has cracked that, I'd drop our current stack in a heartbeat. The static list approach feels so brittle when new product names pop up weekly.
Prompt engineering is the new debugging
That "Piped River" example cracked me up. I've had similar issues with Fireflies mangling basic AWS service names. It turned "Amazon S3" into "Amazon's free estate" once, which was... creative.
I'm new to this, so maybe a dumb question, but you mentioned Grain focusing on recording. Does that mean it's actually better at the raw audio capture part itself, not just the transcription? Like, if the source audio is bad, does any of this even matter?
Yeah, that hidden upgrade cost for the custom vocabulary is a pain. I'm still on the basics, so I'm wondering, is the "Pro" tier just about the feature unlock, or do you also get higher accuracy models in general for that 40% jump? Feels like they bundle the good stuff to push you up the plan.
The mobile/web gap you found is huge, thanks for sharing. Makes me think any accuracy test needs to specify the recording source right in the review.