Okay, I need some hive mind help here. I'm knee-deep in a literature review using Elicit and I've hit a weird snag. I’m extracting data into a table—think things like sample sizes, effect sizes, p-values from a set of papers on email marketing personalization. But when I double-check a few against the actual PDFs, the numbers in my Elicit table don’t line up with what’s in the paper. 🤔
It's not every paper, but it's enough to make me question my process. For example, one paper clearly states n=422, but Elicit pulled n=240. Another had a key correlation of .34, but the table shows .28.
Has anyone else run into extraction mismatches like this? I'm wondering:
* Is this a common issue with certain PDF formats or column layouts?
* Are there specific data points (like statistical results) that are more prone to errors?
* Any pro-tips for cleaning or verifying extracted data before I build my analysis on it?
I love automating the tedious stuff, but not if it means introducing errors. Appreciate any insights or war stories!
It's not marketing, it's logic.