Hey everyone! I've been using SciSpace for a few months now to organize my literature reviews for a university project. Overall, it's super helpful, but I kept running into one annoying thing: sometimes it misses obvious duplicate PDFs, especially if the filenames are slightly different or the metadata is a bit off. My library was getting messy!
I know SciSpace has duplicate detection built-in, but it doesn't catch everything. So, I decided to write a little Python script to clean up after it. I'm still pretty new to data engineering (just learning about ETL and Airflow!), but this felt like a good hands-on project.
The script basically compares documents based on a few things SciSpace might not prioritize: it checks for similar titles using a fuzzy matching library, looks at the publication year, and can even compare the first chunk of text if the PDFs are accessible. It then spits out a list of potential duplicates for me to review manually before I merge anything in SciSpace.
It's not perfect, but it's caught a bunch of duplicates my library missed—like when I had "Smith_2020_Final.pdf" and "smith_2020_published.pdf". I run it as a weekly cleanup job on my local machine. Maybe someday I'll figure out how to schedule it properly with Airflow!
Has anyone else tried something like this? I'd love to know if there are better ways to compare documents or if I'm reinventing the wheel here. Also, if you want to take a look at the code or have suggestions for improvement, let me know!
-- rookie
rookie