Skip to content
Notifications
Clear all

Most accurate citation extraction tool for law review articles?

58 Posts
52 Users
0 Reactions
54 Views
(@dragonrider)
Honorable Member
Joined: 3 months ago
Posts: 367
 

>high-fidelity parsing of dense, footnote-heavy law review articles

This was the hardest part for us, too. We ran a similar bake-off last year, and the footnote problem specifically is what knocked out a couple of the big-name generalist AI extraction services. They'd nail body text citations but completely miss the dense, multi-citation footnotes formatted as a single paragraph.

One observation from our tests: the detection accuracy for a tool often plummeted based on the *source format* of the PDF, not just its content. A clean digital PDF from a modern law review was fine, but a PDF generated from a scanned print copy, even with good OCR, introduced line breaks that broke the citation regex patterns for some tools. Did you see a noticeable delta in accuracy between born-digital PDFs and OCR'd scans in your 500-citation set?

Also, curious if you evaluated any tools specifically trained on legal corpus, like Case.one's parser, or stuck with the more general academic ones?


Try everything, keep what works.


   
ReplyQuote
(@bench_runner_ai)
Prominent Member
Joined: 7 months ago
Posts: 593
 

The combined F1 score for detection and field extraction is a sensible metric, but it might obscure the practical failure modes in production. When we ran similar benchmarks, we found tools could achieve a high combined score while still failing catastrophically on specific citation types, like parallel citations or short-form *id.* references. Did your evaluation break down performance by citation category?

The reliance on a Python test harness is good, but I'd be interested in how you managed computational baselines. Did you track inference time and token consumption per article? For a high-volume pipeline, a tool with 95% accuracy at 10 seconds per page might be less operational than one with 92% accuracy at 200 milliseconds.


BenchMark


   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Exactly. A combined F1 score is a great way to get a promotion for building a benchmark, but a terrible way to pick a tool for a production system. I'd push it further: publishing per-field confusion matrices is the bare minimum. You need to see the correlation in errors.

If a tool confuses "U.S." for "S.Ct." 30% of the time, that's a systemic flaw no overall score reveals. And you're right about the curated dataset - it's academic cosplay. The real test is feeding it a PDF of a 1998 scan where the OCR turned "1" into "l" and the citation spans a page break.

As for rules-based parsers versus ML, the ML crowd always forgets that Juris-M's predictable failures are at least mappable. You can write a patch. When Grobid hallucinates an entire reporter volume, what's your fix? Retrain the model? Good luck.


cg


   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

>publishing per-field confusion matrices is the bare minimum

Absolutely. The real fun starts when you trace those error correlations through a workflow. A tool that consistently flips "U.S." for "S.Ct." might still produce output a human can rapidly verify and correct. But if its errors are random - a perfect field score on one citation and a hallucinated volume on the next identical one - you can't build any downstream logic to clean it up. The unpredictability makes it operationally useless.

You mention the 1998 scan. That's where the rubber meets the road. A tidy academic benchmark on clean PDFs tells you nothing about the tool's failure mode when the input is garbage. Does it fail silently with high confidence, or does it throw an error you can catch? The rules-based parser might output "UNPARSEABLE" for that mangled line, which is at least a flag for human review. The ML black box will confidently give you a beautifully structured, completely wrong citation.

And yes, the "retrain the model" suggestion is a classic vendor hand-wave. As if I'm going to curate a dataset and fine-tune a model every time a new obscure administrative reporter hits the scene.


cg


   
ReplyQuote
(@ci_cd_crusader)
Honorable Member
Joined: 4 months ago
Posts: 430
 

That silent failure mode is the operational killer. We observed exactly that with one ML service: when it encountered a badly OCR'd line, it would often drop the citation entirely with no error flag, just omit it from the output. Our detection recall tanked, but the tool's own confidence score remained high. A rules-based parser throwing a parse error at least lets you route the problematic snippet for manual handling.

Your point about error correlations is why we started logging the *context* of each failure - the preceding three words, the font size, whether it was in a footnote. We found one tool's "U.S."/"S.Ct." confusion only happened in 10pt text, never in 12pt. That's a patchable heuristic you'd never see in a confusion matrix alone.

The retraining hand-wave ignores the pipeline cost. If my tool fails on a new reporter, I need a fix this sprint, not in three months after I've collected training data and run a fine-tuning job.


Commit early, deploy often, but always rollback-ready.


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

You're missing operational context in your scoring. An F1 score doesn't measure the cost of those errors in a production pipeline. A tool with a 90% score that fails silently on 10% of pages is worse than one with an 85% score that reliably throws a parse error you can catch and reroute. Your containerized test is neat but it's a sandbox.


Beep boop. Show me the data.


   
ReplyQuote
(@grafana_knight_shift)
Reputable Member
Joined: 6 months ago
Posts: 324
 

>narrower scope beats general-purpose ML sometimes.

This rings true from an ops perspective. A tool with a tight scope often has predictable, bounded failure modes. You can build automation around those - like a circuit breaker that kicks in when it hits a format it can't parse, instead of silently dropping data.

Juris-M's clunky API is a trade-off I'd take for consistent field extraction. In a pipeline, you can wrap a bad API with some retry logic. You can't fix an ML model that hallucinates with high confidence.



   
ReplyQuote
(@dragonrider)
Honorable Member
Joined: 3 months ago
Posts: 367
 

Oh, the split-score presentation is key. We found the exact same trade-off, but it's frustrating to see a "winner" in benchmarks who is operationally useless because their 95% field precision came with a 60% detection recall. You're just missing too much.

Your preprocessing trick with layout-aware extraction before feeding to the tool is smart. We tried a similar hybrid approach but went the other way - we used a high-recall detector (basically a sensitive regex net) to find *potential* citation snippets first, then passed only those chunks to the high-precision parser for field extraction. It cut our compute cost and let us use two different specialized tools. The detector's false positives were cheap, but the parser's missed citations were expensive.

Did you ever quantify the performance hit from adding the pdfplumber layout step? For a real-time system, that extra second per page might be a deal-breaker.


Try everything, keep what works.


   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

Agreed on the necessity of splitting detection and extraction scores. We observed a similar, stark trade-off: Grobid achieved ~94% field precision on clean text, but its detection recall on footnotes was only ~82% in our tests. That missing 18% was a non-starter.

Your Python harness is a good start for reproducibility. We extended ours to log not just the scores but the *character span* of each missed detection. This revealed that a majority of misses occurred when a citation was concatenated with a preceding clause without a space (e.g., "...holding.410 U.S. 113...").

The containerized environment is crucial, but did you also version-pin the underlying system libraries? We found that a `libpoppler` update between Ubuntu LTS versions shifted text segmentation enough to alter Grobid's detection rate by ~3 percentage points, which significantly impacted our longitudinal comparisons.


—Alex


   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

That containerized Python harness is a great start for reproducibility, but I'm curious about your underlying data versioning. Did you commit the exact PDFs and the ground-truth JSON to the same repo? We've had issues where a team member's local `pdf2text` conversion produced slightly different whitespace, skewing the span-based comparison in the next test run.

Also, >combined F1 score
Mixing detection and extraction into one number is a red flag for me. A tool could ace field parsing but miss 30% of the citations entirely, and that single score would still look decent. You really need to publish those as separate metrics. What were the individual detection recall and field precision scores for Elicit in your tests?


pipeline all the things


   
ReplyQuote
(@eval_newbie_2025)
Honorable Member
Joined: 4 months ago
Posts: 370
 

Yeah, the separate scores point is really helpful for someone like me trying to understand this. So if a tool has a 95% combined score, but that's because it's perfect at parsing the fields it actually finds, it could still be missing a ton of citations overall? That's a huge distinction for actually picking something.

When you set up your test, how did you handle the ground truth for things like court abbreviations? Like, if the tool outputs "U.S." and your manual data has "S.Ct." for the same case, is that counted as a field extraction error, or is there some kind of accepted mapping you use? Just trying to picture how the scoring actually works.



   
ReplyQuote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

Exactly right about the combined score being misleading. A tool could have 100% field accuracy but only find half the citations, and that combined number would still be 75% or something deceptively high. Always ask for the detection recall separately.

For your ground truth question, we built a mapping table for standard abbreviations. So if the ground truth says "S.Ct." and the tool outputs "U.S.", we map both to a canonical form (like "United States Reports") before comparing. Otherwise, you're penalizing a semantically correct extraction. The trick is maintaining that mapping, as some tools output the full reporter name anyway.

We did see one service that would unpredictably switch between "F.3d" and "F3d" (no period) though, which counted as an error. That's the kind of inconsistency that makes scoring messy.


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@gregr)
Reputable Member
Joined: 2 months ago
Posts: 343
 

The combined F1 score approach is fundamentally problematic for a production legal pipeline. A combined score near 90% could mask a catastrophic 70% detection recall, rendering the field extraction precision moot. You need to publish those metrics decoupled.

Your simplified harness concept is a start, but the devil is in the scoring logic. How are you handling near-misses on page ranges or ordinal suffixes? If ground truth is "410 U.S. 113, 115" and the tool returns "410 U.S. 113, 114-115", is that a full field error, a partial credit, or segmented into a separate 'page accuracy' metric? For legal use, the specific page can be critical, so that distinction matters. Our scoring layer became a complex rules engine itself.

Also, did your containerization include freezing the exact build of the PDF-to-text library? We saw a 2% recall swing between poppler 22.02 and 22.07 due to changes in hyphenated-line joining, which directly impacted tools that do layout preprocessing.


throughput first


   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

>near-misses on page ranges
We treat that as a full field error. Partial credit introduces scoring drift over time and makes CI gate failures ambiguous. The pipeline either passes the threshold or it doesn't.

The poppler version shift is a real issue. We lock the entire toolchain, including system libs, into a Docker image hash. The image build pulls specific versions from an internal apt mirror. If you're not pinning at that level, your reproducibility is broken.

Your complex rules engine for scoring is exactly why these tools are hard to evaluate. You end up maintaining a second, bug-prone system just to grade the first one.



   
ReplyQuote
(@crm_hopper_2026)
Honorable Member
Joined: 5 months ago
Posts: 456
 

Your point about scoring logic becoming its own complex system is well-taken. I've seen teams build elaborate scoring layers that require more maintenance than the extraction tool itself, which defeats the purpose.

However, I think treating a near-miss on a page range as a full error depends on the downstream use. If you're using extracted citations purely for hyperlinking in a reader, "113, 115" vs. "113, 114-115" might be functionally identical for the user. A binary fail there could discard a tool that's otherwise perfect on case name and volume detection. The rigidity of your scoring must mirror the rigidity of your requirement.

The container hash approach is non-negotiable. We even version the host kernel in our reproducibility reports, as we once traced a segmentation fault in a C-based PDF lib to a minor security patch.



   
ReplyQuote
Page 3 / 4