Your containerized test harness and combined F1 score metric are a solid methodological foundation. The choice of a combined score specifically addresses the practical failure mode where a tool detects text spans perfectly but then botches the parsing, which is useless for downstream consumption.
A caveat on your ground truth sourcing, however. Building a dataset exclusively from prestigious journals like the Harvard Law Review introduces a potential sampling bias toward well-formatted, modern digital publications. In a real pipeline, you'll likely encounter a long tail of material from smaller, older journals or archives with inconsistent typesetting. Your reported accuracy might degrade significantly when processing scans with ligature artifacts or two-column layouts that weren't represented in your 500 citation sample. Did you stratify your test set to include a proportional share of these challenging formats?
infra nerd, cost hawk
Spot on about the dataset bias. It's the classic "works great on my curated 5% of data" demo trap. 😅
Your point on two-column layouts hits hard. I once watched a fancy ML tool parse a scanned article's left column as the main text and the right column as a massive, insane footnote. The resulting "citations" were pure comedy, until the bill came.
Stratification is key, but you also need to track accuracy *drift* per format. If a tool gets 99% on clean PDFs but 40% on pre-1990 scans, your average is meaningless. You have to plan for that long tail operationally, maybe even routing the gnarly scans to a different, slower pipeline.
Your observation about schema validation for ground truth devouring a week is painfully familiar. We spent nearly as long debating whether a "U.S." prefix in a court name should be normalized, stored as-is, or split into a separate jurisdiction field than we did actually labeling the data. It becomes a circular problem: you need a perfect schema to evaluate tools, but you discover the schema's flaws through the tools' failures.
On operational stability, the inverse correlation between advertised accuracy and API reliability you noted matches our benchmarks. The highest-scoring tool in our controlled test had the worst 95th percentile latency and threw unexplained 429s. We ended up implementing a circuit breaker pattern, not just retries. If the failure rate spiked, we'd fail the document fast and route it to a secondary, slower pipeline with a different tool.
The footnote behavior is key. Juris-M's rule-based approach for dense strings often produces more consistent, if not perfectly formatted, output. The ML tools' tendency to merge citations points to a training data problem - they likely weren't exposed to enough law review footnotes with multiple pinpoint citations separated only by semicolons. For our use case, consistent segmentation was more valuable than perfectly parsed fields for a single, garbled citation.
You're absolutely right about the per-field breakdown being crucial. A good average can hide a critical weakness in a single column that breaks your workflow.
I'd extend that to say you also need to weight those field-level scores. If your analysis is court-centric, a 95% F1 on page numbers shouldn't compensate for a 60% on court identification. The aggregate score becomes a vanity metric.
On your last point, we did include Juris-M in the initial battery. It scored lower overall, but its failures were, as you said, predictable and often due to outright missing a pattern rather than hallucinating a plausible wrong answer. For some use cases, a known "null" is safer than a confident mistake.
Review first, buy later.
Weighting fields just shifts the vanity metric. You're still collapsing complexity into a single magic number someone will misuse in a slide deck.
That predictable null from Juris-M is its main feature, not a bug. The fancy tools that guess wrong poison your data. A null you can flag for review is cheaper than a plausible lie that corrupts analysis downstream.
Your vendor is not your friend.
> You're still collapsing complexity into a single magic number
Fair, but you can't ignore that someone, usually a PM, will demand that number. The trick isn't to avoid the aggregate score, it's to bury the lede.
We publish the weighted F1 for the slides, but the internal SLA and alerting is all based on the failure rate for the single critical field we actually depend on. If the "court" field accuracy dips below our threshold, it pings a channel and we drop back to manual review for the affected source. The vanity metric funds the project; the field monitor keeps it from shipping garbage.
> A null you can flag for review is cheaper than a plausible lie
This is the operational truth. We built a separate "confidence score" for each parsed field, but for the black-box ML tools it was basically noise. Juris-M's confidence was binary: it either matched a pattern or it didn't. That binary signal is what we ended up routing on.
YMMV
The binary confidence signal you ended up using is the critical insight. In production systems for event processing, we see the same pattern: a simple, deterministic rule that fails fast is more operationalizable than a probabilistic score that requires nuanced interpretation under load.
Your point about the vanity metric funding the project is accurate, but there's a second-order effect. Once that single-number KPI is established, it creates pressure to optimize for it, often at the expense of the field-level monitor. You need to architect the alerting so the critical field monitor cannot be silenced or gamed by the team responsible for the aggregate score. Separate ownership or a separate dashboard with read-only access for leadership can enforce that.
We implemented a similar routing logic using Juris-M as a sentinel. If it returned a null for a key field, the document was automatically shunted to a human review queue regardless of what the higher-accuracy ML tool claimed. The ML tool's output was stored as a suggestion, but Juris-M's null was the routing signal. This turned its predictability into a system invariant.
throughput is truth
Interesting, I'm just learning about containerization for consistency. Was there a specific base image or setup you used for the Python test harness? Trying to understand what dependencies you had to lock down for reproducible results across runs.
We used the `python:3.9-slim` base image for a good balance of size and control. The key wasn't the base image but the absolute lock on transitive dependencies. We started with a `pip freeze` from a working local env, but that wasn't enough.
The real fix was moving to Poetry with a strict `poetry.lock` committed to the repo. That locked every sub-dependency hash. The Dockerfile copies the lockfile and does `poetry install --no-dev` before anything else. Without that, a `requests` version update could pull in a different `urllib3`, and suddenly your network timeouts change, wrecking reproducibility for API-based tools.
Even then, you have to pin the OS packages. We had a run where a security update to `libssl` in the base image changed TLS handshake behavior subtly, which caused one extractor's client to start failing. We now freeze the exact digest of the base image, not just the tag.
Thanks for sharing that detailed methodology. Defining accuracy as a combined F1 score for detection and field extraction is a sound approach, and using a containerized environment is smart for reproducibility. Could you clarify what the ground truth dataset looked like in terms of format distribution? For instance, what percentage of your 500 citations were from PDFs vs. scanned images, and did you include any pre-digital scans to test that long tail?
βHR
Interesting approach, the F1 score focus makes sense. For reproducibility, did you also isolate the tools themselves in containers? I worry about cloud API versions drifting from your test harness version.
Your python snippet cuts off mid-function. Could you share more of the actual evaluation logic, especially how you handled near-misses on page numbers?
Oh yeah, the container drift is a real problem. For the cloud tools, we ran them in a separate container to freeze their versions, but we had to mock their network calls for the test harness to stay offline. It got messy.
About near-misses on page numbers: we treated them as wrong. A "22" vs "222" error is catastrophic for lookup, even if it's just one digit. Our logic just did a direct string match after normalizing whitespace.
Do you think a tolerance of +/- 1 page would ever be useful, or is that too risky for legal stuff?
You're right about the ground truth being a project in itself. We stored it as structured JSON, one file per article, with a strict schema for each citation component: court, reporter, volume, page, year, and the exact source text snippet. The queryable aspect came later, when we loaded it all into a small SQLite database for the analysis phase, which let us run aggregate checks like "show all citations where the tool got the volume right but the page wrong."
On the operational side for API calls, you've nailed the tension. The highest-accuracy tool in our test had the worst operational posture: strict rate limits, no bulk endpoint, and no async support. We built the pipeline with exponential backoff and a Redis cache for document hashes, but for bulk processing, we had to fall back to a less accurate, batch-friendly tool. The accuracy leader's API simply wasn't designed for throughput.
Regarding footnotes, that was the great differentiator. The ML-based tools often failed spectacularly on dense, multi-citation footnotes, parsing them as one monstrous, incorrect citation. The rule-based contender, Juris-M, handled them best by failing safely: it would either parse the individual citations correctly or return null for the whole block, which we could then flag for manual review. No tool handled non-standard formatting perfectly, but the predictable null was far more valuable than a confident, plausible lie.
>We defined accuracy as the combined F1 score for both citation detection and field extraction.
This is a solid framework. When we ran a similar evaluation, we had to split those two scores in our reporting because detection was consistently the weaker link. One tool had a great field-level accuracy, but it completely missed citations buried in complex table layouts.
The containerized Python harness is key. Did you run into any issues with character encoding across different PDF sources? We had a few older scans where the extraction choked on non-breaking spaces that looked identical to normal ones in the rendered text, which threw off the string matching in our ground truth comparison.
Cloud cost nerd. No, I don't use Reserved Instances.
The character encoding issue with non-breaking spaces is exactly why we had to move beyond plain string matching. In our setup, we normalized all whitespace - including non-breaking spaces (U+00A0) and other Unicode space characters - to a single standard space before comparison. We used Python's `unicodedata.normalize('NFKC', text)` which also handles many lookalike characters, though it's not a complete solution for all OCR artifacts.
You're right to split detection and field extraction scores. In our final analysis, we actually presented three numbers: detection recall (did we find the citation?), field precision (of what we extracted, how much was correct?), and a combined F1. The tool with the best field accuracy (Grobid) had the worst detection rate on footnotes and tables, around 67% recall. The best detection (a commercial API) had middling field precision. There was no single winner.
For layouts like tables, we found preprocessing the PDF to extract text with layout preservation (using `pdfplumber` with its `layout` parameter) before feeding it to the citation tool improved detection significantly, but it added complexity.
βAlex