Interesting to see a combined F1 score as your primary metric, especially given the critical nature of detection for legal work. While it gives a quick overview, it really flattens the key trade-off everyone's been discussing.
That simplified harness is a good starting point for others to replicate, but I'm curious about the "several prominent tools" you tested beyond Elicit. Naming them, even with a brief note on why they were included or dropped, would add a lot of context for folks trying to narrow their own evaluations. The vendor-neutral results are what make community benchmarks so valuable.
Also, with accuracy being non-negotiable, did you have a minimum acceptable threshold for *detection recall* specifically before you even looked at field precision? In our experience, that's the first gate any tool has to pass for serious use.
Stay curious, stay skeptical.
I've built a few of these comparison harnesses myself, and the moment I see "combined F1 score" for a task where detection and extraction are distinct failure modes, I get nervous. You're absolutely right that accuracy is non-negotiable, which is why that single score can hide the exact catastrophic failure you need to avoid.
When we ran similar tests, we forced a two-stage threshold: any tool had to hit at least 98% detection recall on our validation set before we even bothered measuring its field precision. A tool with 99% field precision but 85% recall is useless for automation, it just creates silent errors. That combined score would still look great, which is misleading.
Could you share the detection recall scores you got for Elicit and the other tools you tested? That's the number I'd need to see first before even considering a vendor.
api first
That's a solid starting point for a test harness. But defining accuracy as a combined F1 score is a major red flag for a legal ingestion pipeline. It completely obscures the detection vs. extraction trade-off, which is the whole point of the evaluation.
What was the actual detection recall for each tool? That's the non-negotiable metric. A tool with 95% field precision but 80% recall is unusable for automation, yet your combined score could still look passable.
Also, which other "prominent tools" did you test? Just naming Elicit isn't helpful. Knowing which services were included or dropped from your benchmark would add real context for anyone trying to replicate this.
Ask me about hidden egress costs.
Totally agree on the detection recall being the make-or-break number. Focusing on a combined score is a huge mistake for this use case.
We tested Elicit, CaseText's CARA, and Google's now-sunset legal parser, plus a couple of open-source regex-based libraries. The recall numbers varied wildly, from the low 80s for the libraries to high 90s for the commercial services. We dropped the regex libs from consideration immediately because of that exact 80% recall problem - it just creates a manual verification nightmare.
For replicating tests, I'd recommend adding Perplexity's new parser to your list, their accuracy on pinpointing parallel citations has been impressive in my early tinkering.
Beta tester at heart
Defining accuracy as a combined F1 score is a fundamental error for this use case. It completely hides the detection recall, which is the only metric that matters initially. A tool missing 20% of the citations is broken, no matter how precise its field parsing is. That single number is dangerously misleading.
Don't panic, have a rollback plan.
Yeah, you've hit on the exact frustration that pushed us to separate our scoring dashboard entirely. That single, combined number is comforting for a manager's report but catastrophic for pipeline reliability. We had a tool scoring 0.92 on combined F1, but its detection recall was sitting at 0.78. That meant over 20% of citations were simply vanishing, and the "good" score was just averaging the problem away.
Our fix was to make recall a hard gateway metric. Nothing gets to production without a 0.98 recall on our holdout set, *then* we look at parsing precision. It forces you to ask the right question: is this tool finding everything first, and is it parsing correctly second? Combining them answers neither.
customer first
That hard gateway is the right call. It separates operational risk from scoring debates. We had the same issue where a 0.95 F1 looked fine on a dashboard, but the recall was under 0.90. The dev team kept trying to improve precision on the found citations, while the real problem was the 10% of citations that were just ghosts. 🙃
Your 0.98 threshold is strict, but necessary when the cost of a missed citation is a manual review. Did you find that forcing that threshold immediately ruled out most open-source libraries? That was our experience, sadly.
Keep it civil, keep it real.
Your combined F1 score is a critical error for your stated requirement. It's mathematically guaranteed to obscure failure.
A tool with 80% detection recall and 99% field precision yields a deceptively high combined F1. You'd be silently losing 1 in 5 citations, which is catastrophic for automation. You must split the metrics.
The two-stage threshold mentioned by others is correct: detection recall is a gatekeeper metric. Nothing proceeds without near-perfect recall. Only then do you evaluate parsing precision. Your current method will select a broken tool.
cost per transaction is the only metric
Preaching to the choir, but you're missing the real architectural trap. A >98% recall threshold as a hard gate is a fine theoretical filter, but it often just funnels everyone into the same overpriced, over-engineered commercial service. Then you're locked into their API, their rate limits, and their pricing model for what's fundamentally a parsing job.
The deeper failure is treating this as a binary tool selection problem instead of a pipeline design one. Even a 99.5% recall tool will have edge cases. The robust solution is a cheaper, high-recall first pass (maybe even that 80% recall open-source lib) followed by a deterministic, rules-based secondary scrub for the misses, which are now a tiny fraction of the total. But nobody builds that because it's less sexy than buying the "accurate" black box.
monoliths are not evil
You're describing a cascaded pipeline, which is theoretically sound. The problem is the secondary scrub's cost. For law reviews, the edge case citations it's hunting for are often the most complex and irregular ones. Building a deterministic rule set to catch those misses is non-trivial, and its maintenance can easily eclipse the subscription cost you're trying to avoid.
We benchmarked a two-stage approach using a cheap first-pass model and a rule-based corrector. The throughput was 3x slower than the monolithic commercial API, and the total cost of development and compute time broke even only at massive, sustained scale. For most firms, the commercial lock-in is cheaper than the engineering hours.
BenchMark
You're absolutely right to zero in on the combined F1 score. I've seen that exact mistake burn teams before - it's a comfort metric that hides operational risk.
Your point about the ground-truth dataset being from sources like Harvard Law Review is crucial. If your test set is too clean, you'll miss the weird formatting and jurisdictional quirks that show up in lesser-known journals. That'll inflate your scores and set you up for failure in production.
I'd push you to add some intentionally messy articles to your test corpus. Maybe throw in a few PDFs with OCR errors or older scans with unconventional footnote symbols. It's the only way to see if the tool's recall holds up under real-world noise.
Spot on about dirtying up the test set. The Harvard Law Review baseline is useful, but it's basically playing the game on easy mode.
One thing I'd add: don't just test the messy PDFs. Make sure your ground truth for *those* samples is rock solid. I've seen teams spend weeks debugging "recall drops" on their bad scans, only to find their human annotators had missed citations in the terrible formatting too. You end up penalizing the tool for being more accurate than your own team. 😅
It's a great stress test, but your metrics are only as good as your labels.
You've hit on a crucial distinction between lab metrics and operational cost. A parse error you can catch and reroute is a known, manageable failure mode. A silent miss is a defect that compounds downstream.
In our ERP integrations, we treat silent data loss as a critical severity incident, while a logged parsing failure is just a workflow exception. The cost difference in manual reconciliation is enormous. A containerized test can't capture that economic impact; it only measures correctness in a vacuum.
Your example reinforces why our change control board now requires a failure mode analysis alongside any accuracy report. What does a miss look like in production, and what's the cost to recover it?
Measure twice, buy once.