Skip to content
Notifications
Clear all

Showcase: My script that cross-checks ChatPDF outputs with source PDFs.

28 Posts
28 Users
0 Reactions
94 Views
(@hannahr)
Reputable Member
Joined: 3 months ago
Posts: 285
Topic starter   [#22117]

After our team migrated a massive archive of legacy contracts to a new vendor portal, we learned a hard lesson: automated extraction tools like ChatPDF are incredibly useful, but you can't fully trust them without a verification step. We had a few close calls where subtle but critical clauses were misrepresented.

To mitigate this risk, I built a simple Python script that cross-chexes ChatPDF's answers against the actual source PDF text. It doesn't rely on perfect OCR; it just checks if the key phrases and terms from the answer genuinely appear in the document's raw text. This has been invaluable for our compliance audits.

Here's the core workflow my script automates:

* **Input:** You provide the PDF file path and paste the ChatPDF answer.
* **Processing:** It extracts text from the PDF (using `PyPDF2` for simplicity) and preprocesses both the source text and the answer (lowercases, removes punctuation).
* **Verification:** It checks for the presence of significant nouns/terms from the answer within the source text. It ignores common stop words.
* **Output:** It generates a report showing:
* A confidence score (percentage of key terms found).
* A list of any claimed "key terms" not found in the source.
* The relevant source text snippets for verification.

For example, if ChatPDF claims a contract includes a "liquidated damages clause of 15%," the script will flag it if "liquidated damages" and "15%" do not co-occur in the actual PDF. It's caught several instances where ChatPDF confidently conflated details from different document sections.

This isn't a silver bullet for hallucinations, but it adds a crucial, repeatable checkpoint. It's especially helpful when dealing with large document batches where manual spot-checking is impossible. I'm happy to share the core logic if anyone is building similar safeguards into their workflow.

- h


Data is sacred.


   
Quote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Oh, this is rich. So after ChatPDF gives you a wrong answer, you need a second script to verify if the first script's work is correct? That's not a solution, that's a symptom of the problem.

You're checking if "key phrases" appear in the raw text, but what if ChatPDF hallucinates a *logical connection* that your script can't catch? For instance, it correctly pulls "Party A" and "liable" and "force majeure" from the text, but then tells you Party A is liable under force majeure when the clause actually says the opposite. Your verification score would be high, but the answer is dangerously wrong.

You've just automated the process of looking for the words, not the meaning. I'm curious if your compliance team knows they're just getting a glorified keyword matcher.


cg


   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 3 months ago
Posts: 380
 

Your point about logical connection hallucination is valid and gets to the core limitation of any keyword-based verification. This is precisely why the script should be viewed as a first-pass filter, not a semantic validator.

A practical next step would be to extend the script to also check for negation phrases around the matched terms. For a simple implementation, you could have a list of negation triggers, like "not liable" or "exempt from", and then examine a window of text around each matched key term from the ChatPDF answer. The report could then flag potential contradictions.

In complex contracts, even that isn't sufficient, but it would elevate the check from a pure keyword matcher to a basic contextual scanner. The real goal is to reduce the manual review load to only those answers that pass the initial term check but might contain inverted logic, which is a smaller subset.


null


   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

That's a smart approach! I've seen similar needs in our gitops flows where we auto-generate K8s manifests from docs. One thing I'd add: consider using `pdfplumber` instead of `PyPDF2` for extraction, it handles tables and formatting way better, which can be critical for contract data.

Have you thought about adding this as a pre-commit hook or a GitHub Action? You could run it automatically when PDFs are added to a repo, then attach the verification report to a pull request. That'd make your compliance audit trail part of the code review process. 😄


git push and pray


   
ReplyQuote
(@finnm)
Reputable Member
Joined: 3 months ago
Posts: 280
 

Oh, I just started trying to automate some of our vendor onboarding stuff with ChatPDF. The confidence score report you mentioned is a great idea.

Do you have a rough threshold for what score is "good enough" to trust? Like, if it's over 80%, do you still manually check everything?

Thanks for sharing this, it's a big help.



   
ReplyQuote
(@fionaj)
Estimable Member
Joined: 2 months ago
Posts: 203
 

Great question about the threshold! I'd be really careful about setting a fixed number like 80%. In my small tests, a high score just means the words are there, not that the answer's logic is correct - which is what user980 mentioned above.

Even with a 95% score, I'd probably still skim the answer against the PDF if it was about money or deadlines. For something minor, maybe that's when you trust it.

How are you planning to test what a good threshold is for your onboarding docs?



   
ReplyQuote
(@chrisw2)
Reputable Member
Joined: 2 months ago
Posts: 309
 

Exactly. A fixed threshold is a trap. The "good enough" score depends entirely on what you're asking about.

I set up alerts for my script based on the answer's *category*. Questions about contract dates or monetary values get flagged for review at any score, even 100%. For something like "Does this document mention GDPR?", a high keyword score is probably fine.

You should test by creating a validation set of Q/A pairs you know are correct and a few where ChatPDF hallucinates. Plot the scores. You'll likely see two overlapping clouds, which tells you a single threshold won't work.


Run it yourself.


   
ReplyQuote
(@gregoryt)
Reputable Member
Joined: 2 months ago
Posts: 418
 

That's a smart idea, checking for negation phrases. I was trying to think of a simple way to catch some of that context.

Wouldn't a list of triggers get pretty long and specific to different document types? Like "exempt from" for contracts, but "not applicable" or "failed to" for audit reports.

Maybe you could also flag sentences where ChatPDF's answer uses a strong, definitive word like "must" or "always," and then see if the source text around the matched terms uses weaker language like "may" or "often." That feels like another common way meaning gets flipped.



   
ReplyQuote
(@ci_cd_enthusiast)
Honorable Member
Joined: 7 months ago
Posts: 382
 

You're spot on about the trigger lists getting unwieldy. That's the main reason I moved away from maintaining them manually in my own pipelines.

Your point about flagging definitive vs. weak language is excellent - it's another layer of context. I ended up using a lightweight sentiment/scoring library (like `textblob` for a quick PoC) to compare the "certainty" tone between the answer and the source snippet. If ChatPDF says "the vendor *must* indemnify" but the source text says "the vendor *may* be required to indemnify," that gets flagged as a high-priority mismatch, even if all the keywords are present.

It's still not perfect for complex logic, but layering these checks - keywords, negation, tone - has caught a lot of the low-hanging fruit before our legal team even sees the report.


Pipeline Pilot


   
ReplyQuote
(@gracem)
Reputable Member
Joined: 3 months ago
Posts: 294
 

That's the whole point of the script - to flag those exact cases. If ChatPDF answers "Party A is liable" and my extractor finds the clause says "Party A is *not* liable," the confidence score plummets. It's a sanity check, not a full semantic analysis.

Your compliance team analogy is a bit off. It's more like giving the intern a highlighter and a checklist before the lawyer reads the document. It catches the obvious errors fast, so they can focus on the tricky logical connections.


Automate everything.


   
ReplyQuote
(@darrenk)
Honorable Member
Joined: 3 months ago
Posts: 392
 

Exactly the kind of sanity check we need. I'd suggest adding a step to track *where* in the doc the matches happen. If all the key terms for an answer are scattered across 50 pages, it might be stitching ideas together incorrectly. Seeing the page numbers in the report would make that obvious.


dk


   
ReplyQuote
(@henryj)
Reputable Member
Joined: 2 months ago
Posts: 224
 

You're ignoring the biggest cost - processing time and infrastructure. You think your legal team is going to run this manually every time? Or are you now standing up a server to batch-process this archive? That's vendor lock-in of a different sort. You've traded ChatPDF's potential inaccuracy for your own operational overhead.


Show me the data


   
ReplyQuote
(@davidw)
Reputable Member
Joined: 3 months ago
Posts: 320
 

PyPDF2 for text extraction is asking for trouble with anything but the simplest PDFs. It's a toy library. You'll miss half your content if the contracts have tables, forms, or scanned pages.

This basic keyword check also ignores logical operators. The phrase "must indemnify" checking out is useless if the clause is "must indemnify unless otherwise agreed in Schedule B." Your confidence score would be 100%. That's not a sanity check, it's a false sense of security.


Trust but verify.


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

That's a solid start, and I love seeing folks build these kinds of safety nets. Been there with invoice extraction myself, had an AI swear we owed a million dollars... it was for a million *yen*.

PyPDF2 will definitely leave you hanging if your contracts have any scanned signatures or tables, though. For a quick upgrade without much fuss, try `pdfplumber` instead. It's still pure Python, but it does a much better job pulling text from trickier layouts.

Your confidence score is the right first step, but like others said, it's just step one. I'd throw in a simple check for negation words ("not," "never," "exempt") near your matched keywords. It catches the most dangerous flips in meaning with almost no extra code.


it worked on my machine


   
ReplyQuote
(@integration_maven)
Reputable Member
Joined: 6 months ago
Posts: 261
 

This is a great practical approach. Moving from "trust the output" to "measure the output's alignment" is the key mindset shift.

Your point about it being a sanity check rather than full semantic analysis is crucial. It's a first-pass filter. The real value I've found isn't just in the score, but in the report it generates for a human reviewer. That list of "missing key terms" immediately directs their attention to what might have been fabricated or misassociated.

My only immediate tweak would be to add a length check for the answer. If ChatPDF produces a three-paragraph summary but your script only finds three matching keywords scattered in a 50-page document, that's a different risk profile than a one-sentence answer with the same score. The context gets thinner as the answer grows.


IntegrationWizard


   
ReplyQuote
Page 1 / 2