Skip to content
Notifications
Clear all

Showcase: My script that cross-checks ChatPDF outputs with source PDFs.

28 Posts
28 Users
0 Reactions
95 Views
 danw
(@danw)
Reputable Member
Joined: 3 months ago
Posts: 387
 

PyPDF2 will fail on any real contract with scanned pages or tables. Swap it for pdfplumber or a dedicated service. Your confidence score is meaningless if it's built on partial text.

The list of missing terms is useful, but you need a proximity check. Finding "indemnify" on page 1 and "Party A" on page 20 doesn't mean ChatPDF is right.



   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

Thanks for pointing out pdfplumber, I'll swap that in next. I hadn't even thought about scanned pages breaking everything.

>need a proximity check
That's a really good call. My test cases were short, so I didn't see that problem yet. How do you decide what's close enough in a long document? A certain number of words? Or looking for the same paragraph?



   
ReplyQuote
(@charlotteb)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Glad you're switching to pdfplumber, that'll save you a world of pain. On the proximity check, you've hit on the real challenge. A fixed word distance will fail you, because a contract might list parties in a separate definitions section. I've had better luck with a two-tiered approach.

First, check if the terms appear within the same logical block - like a numbered clause or a distinct paragraph. That's your strongest signal. If they're in different sections, then fall back to a broader window, say within 2-3 pages, and flag it for manual review. This catches the "Party A on page 1, indemnify on page 20" problem without demanding they be adjacent sentences.

The real trick is making that proximity threshold configurable per document type. A 10-page research paper is different from a 200-page merger agreement.



   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

The two-tiered proximity check is a great framework, it's exactly the kind of practical heuristic that makes these tools useful. Your point about configurable thresholds per document type is crucial.

For marketing and sales contracts, I've found a third layer helpful: checking proximity within the same *heading*. If "Termination" and "Fees" both appear under a "Payment Schedule" sub-clause, that's a stronger link than just being on the same page, even if the clause spans half a page itself. pdfplumber's layout detection can sometimes pick those headings up.

But you're right, the definitions section is the classic trap. My script once flagged a key term as "missing" because the answer used "Client" and the PDF only had "The Company" - but they were defined as the same entity on page two. That's a semantic layer this kind of check won't ever solve.


Happy testing!


   
ReplyQuote
(@henryf)
Reputable Member
Joined: 3 months ago
Posts: 291
 

Good point about definitive vs. weak language. That's another nuance the keyword check would miss entirely.

Building that into a script would be tricky. You'd need to parse sentence structure in both the answer and the source, not just find words.

Maybe a simpler step is flagging answers containing strong modals (must, shall, will) for extra scrutiny, even if keywords match. It forces a human to look.



   
ReplyQuote
(@data_diver_dan)
Honorable Member
Joined: 6 months ago
Posts: 455
 

That's a good practical filter. Flagging answers with strong modals for review creates a high-priority queue without needing full NLP.

The caveat is that legal documents use those modals precisely. A match on "shall pay" within the correct clause is what you *want* to see. The risk is when the model hallucinates a new obligation, attaching "shall" to a keyword that exists elsewhere. So the flag's utility depends heavily on the preceding keyword and proximity checks.

You could extend it by also flagging weak modals like "may" or "could" when the source contains a strong one. That's often where ChatPDF softens a firm requirement, which is just as dangerous.


Garbage in, garbage out.


   
ReplyQuote
(@emilyf)
Reputable Member
Joined: 3 months ago
Posts: 227
 

This sounds incredibly useful. How do you define "significant nouns/terms" for your verification step? Is it just keywords over a certain length, or are you using something like NER to pick them out?

Also, have you found a sweet spot for your confidence score? Like, anything above 80% is probably safe, or does it vary too much by document type?



   
ReplyQuote
(@brian)
Reputable Member
Joined: 3 months ago
Posts: 282
 

The problem is your confidence score is based on "significant nouns/terms" but you never define how you get them. That's the whole foundation.

If you're just splitting the answer by spaces and ignoring stop words, you're going to flag answers as risky just because they contain pronouns or synonyms. The terms need to be contextually significant, not just long words.

Before you tweak the scoring, fix how you pull the terms. Otherwise the number is meaningless.


Trust but verify.


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 5 months ago
Posts: 404
 

Exactly. That's the core of the garbage-in-garbage-out problem. If your term extraction is naive, your score is just noise.

I ran into this with vendor contracts. Pulling nouns gave me "Company," "Agreement," and "Effective Date" on every check, drowning out the real signal like "Liquidated Damages" or "Termination for Convenience."

You need a domain-specific stoplist at a minimum. Stripping out generic legal boilerplate nouns is the first filter. After that, I look for terms that appear in the source PDF less than, say, ten times. A rare term in the answer that's also rare in the source is a much stronger anchor than a common one.


Cloud costs are not destiny.


   
ReplyQuote
(@calebw)
Reputable Member
Joined: 2 months ago
Posts: 233
 

You've nailed the critical filter. A domain-specific stoplist is non-negotiable, and the 'rare term' heuristic is clever. It reminds me of TF-IDF but for validation.

The danger is over-trimming. In some technical docs, "Company" is the key actor in a critical sentence, and stripping it could leave you with nothing but verbs. I've seen scripts fail because they pruned all the 'boilerplate' only to miss that the hallucination was about *what* the Company supposedly agreed to.

A hybrid approach works better: use the stoplist and rarity score, but also preserve a shortlist of core document entities pulled from a title page or definitions section. Those common-but-critical terms get a pass if they're in the right proximity to your rare anchors.


It's just pattern matching


   
ReplyQuote
(@alexb)
Reputable Member
Joined: 2 months ago
Posts: 257
 

Great point on the hybrid approach. I've been burned by stripping out "Client" only to have the hallucination be about the client's obligations, leaving no trace left to check.

That shortlist of core entities is key. For marketing contracts, I automatically pull the first 3-5 capitalized nouns from the first two pages as my "core actors" and exclude them from the stoplist. It's crude, but it prevents that over-trimming.

The proximity between a rare anchor and a core actor then becomes the real confidence signal.


Data > opinions


   
ReplyQuote
(@ashp99)
Honorable Member
Joined: 3 months ago
Posts: 377
 

Great question. Honestly, I don't trust a raw threshold for anything mission-critical. The score is more of a prioritization flag for me.

>if it's over 80%

For vendor contracts, an 80% on a clause about deliverables might be fine. But an 80% on payment terms or liability? I'm reading every word. The document type and clause severity weigh more than the score.

My rule is: anything under 70% gets a full manual review, 70-90% gets a spot-check focused on the flagged weak links, and over 90% I'll skim. Even a high score can miss a subtle but critical swap of "may" for "shall."


data over opinions


   
ReplyQuote
(@crm_hopper_alt)
Reputable Member
Joined: 4 months ago
Posts: 357
 

A confidence score based on "significant nouns/terms" from `PyPDF2` text? You're building on a shaky foundation right out of the gate.

PyPDF2's text extraction is notoriously unreliable with complex layouts. Your script might be checking terms against a source text that's already scrambled - missing headers, jumbling columns, skipping pages. So your score could be low because ChatPDF is wrong, or because PyPDF2 failed to pull the right clause. How do you know which?

You need to validate your text extractor first, or at least log extraction warnings. Otherwise you're just layering one uncertainty on top of another.


been there, migrated that


   
ReplyQuote
Page 2 / 2