Your custom script is the right track. But PyPDF2 is the weak link. Switch to pdfplumber. Set `layout=False`. It'll handle the diacritics better, assuming your PDFs are scanned images of text and not actual text layers.
Your bigger problem is trusting any off-the-shelf summarizer with nuance. They're all trained to strip it out.
Even with a clean extract, you'll need to budget for that local LLM cost. It's not free. You're paying in hardware wear and time. Everyone forgets to factor that in until their laptop fan screams for an hour.
Read the contract
Yeah, pdfplumber over PyPDF2 is a solid move for extraction. Layout=False was a game changer for me on some old scans.
You mentioned piping to a local LLM - have you played around with specifying the encoding explicitly in your script? Sometimes the default 'utf-8' doesn't cut it for older docs, and you need to force 'utf-8-sig' or even 'latin-1' before the text hits the model. Saved me a few times.
The "sterile" output part is a whole other battle though. Even with perfect text, getting an LLM to care about rhetorical style... good luck 😅
Pipeline Pilot
You've just described my exact problem. I'm also working with texts that have a lot of diacritics and the corruption is a nightmare.
That careful prompting idea from the later posts, to tell the LLM to act like a research assistant in your field, seems like the only way to get past the "sterile" output. But I'm still stuck on the first step of reliably getting the clean text out of the PDF. How do you validate your extraction is actually clean before you waste time running a summary? Is it just manual spot-checking?
Manual spot-checking feels unsustainable. Could you build a quick validation step into your script? Something that samples random lines from the extracted text and flags any that don't match a regex for common diacritics in your language.
That research assistant prompt is key for the output, but you're right, it's pointless if the input is garbage. I started piping a cleaned sample through a small, fast model first, just to see if it chokes on the characters before running the full summary. It adds a step, but saves more time.