Skip to content
Notifications
Clear all

Check out my accuracy test on legal clauses - surprising failures.

2 Posts
2 Users
0 Reactions
4 Views
(@emmaj)
Estimable Member
Joined: 1 week ago
Posts: 92
Topic starter   [#14463]

Hey everyone, I've been deep in the weeds testing various AI tools for parsing complex documents, a big part of my MarOps workflow for compliance and contract review. Naturally, I gave ChatPDF a thorough run with some dense legal clauses.

I was optimistic, but my accuracy check revealed some pretty surprising and specific failure modes. I uploaded a standard SaaS MSA and asked it to summarize key obligations, liabilities, and termination clauses.

Here’s what I found:
* **Nuance Loss on Limitations of Liability:** It correctly identified the liability cap clause but completely missed the **carve-outs** for confidentiality breaches and indemnification. This is a critical distinction that changes the entire risk profile.
* **"Automatic" Renewal Misinterpretation:** It flagged a clause as "automatic renewal" when the language clearly stated renewal required **positive written consent** from both parties 30 days prior. This is a major error for subscription businesses.
* **Conflicting Obligations Not Flagged:** In the data protection section, it listed obligations from both parties separately but didn't highlight a potential conflict in specified notification timelines. A human (or a good checklist!) would spot this instantly.

I used a **ground truth checklist** I've built for vendor contracts to verify each point. It seems ChatPDF is great for surface-level extraction but struggles with **conditional logic and cross-referencing** within a document.

Has anyone else run into similar issues with legal or highly technical PDFs? I'm curious if there are specific prompt techniques or pre-processing steps that improve accuracy for this use case. I was hoping to integrate something like this for a first-pass analysis, but these failures are a bit too risky for my comfort.

Cheers!



   
Quote
(@henry)
Estimable Member
Joined: 1 week ago
Posts: 79
 

That nuance loss on liability carve-outs is a huge red flag. I've seen the same issue when using AI summaries for vendor risk assessments - it glosses over the very exceptions that make a contract risky or safe.

Have you tested its performance on amendments or redlined versions? I've found that's where most tools, not just ChatPDF, really fall apart. They struggle to track changes and often present the original clause as the current one.


Cheers, Henry


   
ReplyQuote