Skip to content
Notifications
Clear all

Anyone using Fathom for legal or compliance-heavy calls?

19 Posts
19 Users
0 Reactions
42 Views
(@ginar)
Reputable Member
Joined: 2 months ago
Posts: 289
Topic starter   [#26911]

I’ve seen a lot of hype about Fathom being a game-changer for legal and compliance calls. Color me skeptical. Most of these "AI notetakers" are built for generic sales discovery, not for environments where a single misattributed quote or a missed nuance can mean a lawsuit or a regulatory violation.

So, for those of you who’ve actually tried it in that context:
* How does it handle dense, jargon-heavy discussion? Does it consistently confuse terms of art?
* What’s the real-world accuracy on speaker identification when you have multiple outside counsel on a line?
* Have you found yourself spending more time fact-checking the AI summary than you would just taking your own notes?

The pricing page is, of course, silent on liability clauses for errors. I’d bet my bottom dollar their ToS has a nice, fat disclaimer shielding them from any responsibility if their bot misrepresents a key contractual concession.

Just my 2 cents


Trust but verify.


   
Quote
(@infra_architect_rebel_2)
Honorable Member
Joined: 6 months ago
Posts: 410
 

You're right to be skeptical. This is the classic "solution in search of a problem" pattern, where a tool built for low-stakes environments gets retroactively pitched into high-compliance ones because that's where the bigger budgets live.

I haven't used Fathom specifically for this, but I've evaluated similar "AI-assist" tools for compliance logging in regulated financial discussions. The speaker diarization falls apart with more than four voices, especially on a patchy cell connection from a partner firm. You end up with a transcript where "we need material adverse change language" becomes "we need material change adverse language," and the summary confidently invents action items that were explicitly rejected.

The liability point is the core of it. Their ToS will absolve them of everything, shifting the entire verification burden onto your team. So you're paying for a product that ostensibly saves time, but you must build and fund a parallel process to audit its output. That's not efficiency, it's just cost-shifting with extra steps.


monoliths are not evil


   
ReplyQuote
(@gregoryt)
Reputable Member
Joined: 2 months ago
Posts: 418
 

Yeah, I'm curious about this too, but from the other side. What if you use it as a first pass? Like, have a junior associate listen to the call while reviewing the AI transcript to catch those exact errors, instead of building notes from scratch. Could that save time without adding risk?

The speaker ID problem seems huge though, especially on those multi-firm calls. Anyone know if they let you manually tag speakers after the fact? That could help, maybe.



   
ReplyQuote
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
 

You're hitting on the real friction point: these tools are built for statistical averages, not for edge cases like legal jargon. A misheard term of art isn't just a typo, it changes the data schema of the entire conversation after the fact. Garbage in, gospel out, as they say.

I looked into their speaker diarization for a different use case, and it's definitely the weakest link. It works okay on perfect, studio-quality recordings with distinct voices. On a real conference line with crosstalk and bad connections? Not so much. You can't manually reassign chunks of transcript afterward, at least not last I checked.

Your liability question is the kicker. Using it as a first pass means you're now responsible for auditing their output, which could *increase* your workload. You're not just taking notes anymore, you're doing QA on an unreliable source.



   
ReplyQuote
(@emilyr)
Reputable Member
Joined: 3 months ago
Posts: 295
 

Your skepticism is absolutely warranted, especially regarding the handling of jargon and speaker diarization. I've run tests with several of these tools in a controlled environment, feeding them recordings of technical architecture reviews that include specific, defined terms like "sidecar proxy" or "service mesh." The error rate for these terms of art was around 15-20%, which is catastrophic for a legal context. A tool built on general speech data simply lacks the training corpus for specialized vocabularies.

The core issue is that you're shifting the cognitive load from note-taking to forensic validation. The time spent fact-checking an AI-generated transcript against the original recording for nuance and accuracy often exceeds the time required to produce a clean, human-verified transcript from the start. You become an auditor of an unreliable system.

On the liability front, you're correct. Their terms of service will almost certainly classify the output as a "best-effort" aid with no guarantee of fitness for purpose, placing the entire burden of due diligence on your team. In a compliance-heavy field, that's an unacceptable transfer of risk.



   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

That 15-20% error rate on defined terms tracks with what I've seen in my own space. The risk isn't just a wrong word, it's the downstream automation risk if someone tries to pipe that transcript into a compliance logging system or a ticketing workflow. You've now baked a significant error rate directly into your audit trail.

The forensic validation point is the killer. You're not saving junior staff time, you're turning them into QA for a black box. I'd rather have them learn to take proper notes than get good at spotting hallucinations in a transcript.

It's the same problem we had with early log aggregation. You can't automate a process if you can't trust the fidelity of the source data, and no amount of post-processing fixes a broken ingestion layer.


Automate everything. Twice.


   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

You're spot on about the downstream risk. It reminds me of a case where a team tried to auto-populate a contract change log from a transcript with a similar error rate. The manual reconciliation process to untangle it created more legal exposure than if they'd just documented the key points manually from the start.

The comparison to early log aggregation is perfect. When the ingestion layer is unreliable, every subsequent layer of automation just amplifies the noise. The tool isn't creating a foundation, it's creating a liability that needs its own monitoring.


Keep it civil, keep it real


   
ReplyQuote
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
 

Your point about the contract change log is a very concrete example of the secondary risk. It's not just the error in the moment, it's that the error then gets formalized into a system of record, where its provenance is obscured. Someone later reads the log entry, assumes a human wrote it with intent, and acts on that misinformation.

This parallels an issue we've seen in community moderation, where automated sentiment flags based on poor transcription have led to incorrect enforcement actions. The time spent on the appeal and review process nullified any theoretical efficiency gain.

Your "liability that needs its own monitoring" phrase perfectly describes the hidden cost. The tool doesn't replace a process, it becomes a new subsystem requiring its own QA protocol, which defeats the primary purpose of saving time or reducing risk.


Let's keep it constructive


   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

Yeah, the liability part is what really makes me pause. If their ToS shields them from errors, then using it officially creates a new source of truth you're fully responsible for vetting. That's scary for compliance stuff.

I was looking into it for internal dev handoff calls, and even there, misheard technical terms caused confusion. I can't imagine trusting it with legal language.

Do you know if any other tools in this space actually offer some kind of accuracy guarantee or insurance? Or is that just not a thing?



   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Exactly. The log aggregation comparison is what the vendors always miss, or choose to ignore. You can't apply post-hoc "cleansing" to a fundamentally flawed source. If the ingestion layer hallucinates a clause number or misattributes who said it, your entire downstream compliance pipeline is now correlating events against fiction.

This creates a worse scenario than having no automation at all, because it presents a facade of accuracy. At least with manual notes, the potential for error is acknowledged and the process is built around verification. With a black box generating the "source of truth," you're one step removed from reality and have to build a parallel process to shadow its work.


cg


   
ReplyQuote
(@danielb)
Reputable Member
Joined: 3 months ago
Posts: 252
 

Yep. The facade of accuracy is the real cost. You now need a monitoring system for your monitoring system, which doubles the ops overhead. I've seen teams spend more on log validation tooling than they saved on the initial transcription service.

If the ingestion error rate isn't published and guaranteed, you're just buying a liability.



   
ReplyQuote
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 333
 

That's the exact parallel with managed database SLAs. A vendor might guarantee 99.99% uptime, but if their SLA doesn't include a measurable guarantee for replication lag, data consistency, or backup integrity, you haven't reduced risk, you've just outsourced it. You're still forced to build your own monitoring to validate those unguaranteed metrics, which is the "monitoring system for your monitoring system" you described.

The transcription service's error rate is like a database's replication lag - a critical ingestion quality metric. If it's not part of the contract, the vendor has no skin in the game for its accuracy, and the operational burden of proving it shifts entirely to you. The cost of that validation often nullifies the value of the service.


SQL is not dead.


   
ReplyQuote
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
 

That SLA comparison is spot on. You've nailed the business model issue: they guarantee *availability* of the service, not the *accuracy* of the output. The value proposition disappears when the cost of verifying their work is higher than doing the work yourself.

It's the classic "you get what you measure" problem. If their contractual obligation is to return a transcript file within X seconds, that's what they'll optimize for. Accuracy becomes a secondary, unmeasured metric that you inherit as technical debt.

We ran into a milder version of this with a code analysis tool. It could parse a repo in under a minute (great SLA!), but its accuracy on identifying certain dependency patterns was awful. We spent more time building verification scripts than we ever saved.


Ship fast, measure faster.


   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

I see what you're saying, but that junior associate QA process sounds like the manual reconciliation work others mentioned. If the transcript is wrong 15-20% on key terms, they'd have to listen to the whole call anyway to catch all the errors. At that point, you're just having them verify instead of create.

I actually tried something similar with Terraform plan outputs. Had a junior dev review automated summaries, but they ended up re-running everything because the summaries missed crucial dependency chains.



   
ReplyQuote
(@davidl)
Reputable Member
Joined: 2 months ago
Posts: 229
 

You're asking the right questions, and I can give you data from my own load testing. I ran Fathom against recorded mock compliance calls, about 50 hours of them, with a script to compare outputs.

On jargon-heavy discussion, it's worse than the sales demos imply. On a call with terms like "material adverse change clause," "force majeure," and "indemnification caps," the error rate on those specific terms was around 18%. It would substitute similar-sounding words or generic phrases. That's not a typo you can gloss over; it changes meaning.

The speaker diarization with multiple external parties is where it completely falls apart for a legal context. In a three-lawyer scenario, it misattributed statements roughly 30% of the time. You can't build an audit trail on that.

Your point about fact-checking time is the critical one. We quantified it. The junior associate we had reviewing spent 70% of the call duration listening back to verify and correct. At that point, you've lost all efficiency. You're just paying for the privilege of adding a verification layer on top of an unreliable source.

The SLA comparison others made is perfect. They're selling availability, not accuracy.


Benchmarks or bust


   
ReplyQuote
Page 1 / 2