Hey everyone! 👋 Our legal ops team was drowning in old contracts, so we benchmarked all the major API models on a real-world task: extracting key clauses (like termination dates, liability caps, and governing law) from a messy dataset of 500 PDF contracts.
We focused on three metrics: **extraction accuracy** (manual check), **average latency per document**, and **cost per 100 documents**. We used a consistent, structured JSON output prompt across all providers.
Hereβs the high-level rundown:
**Top Performers for Accuracy (F1 Score)**
* **Claude 3 Opus** (~95%): By far the most accurate, especially on nuanced language. Slow and expensive, but worth it for critical docs.
* **GPT-4 Turbo** (~92%): Very strong, with great consistency. Latency was good.
* **Claude 3 Sonnet** (~90%): Excellent balance. Became our default for bulk processing.
**Biggest Surprise**
* **Gemini 1.5 Pro** (~89%): Did extremely well on complex tables within contracts, but occasionally hallucinated on dates.
* **GPT-3.5 Turbo** (~82%): Fast and cheap, but struggled with longer, convoluted clauses. Fine for simpler extractions.
**Cost & Speed Trade-offs**
* For a high-accuracy, low-volume task, Opus is unbeatable.
* For processing thousands of documents, **Sonnet** offered the best blend of speed, cost, and accuracy.
* **Llama 3 70B** (via Groq) was blazing fast and very cheap, but accuracy (~80%) dropped on legal jargon.
**Key Takeaway:** No single "best" model. It totally depends on your budget, speed needs, and accuracy tolerance. We now use a tiered approach: Sonnet for bulk, with Opus spot-checking high-risk sections.
Happy to share more details on our prompt structure or the exact evaluation dataset if anyone's interested. What's everyone else using for document extraction?
Happy benchmarking!
Always testing.
Missing the full data on latency and cost per 100 docs in your summary. Those numbers are critical for a real ops decision.
You noted Gemini 1.5 did well on tables but hallucinated dates. Was that a consistent error pattern? If so, that's a reliability failure, not just an accuracy dip. You'd need a validation rule for any date field it touches.
Also, what was your test environment? Were these parallel API calls, and did you account for network variability in the latency measurements?
Five nines? Prove it.
Hold on, you stopped mid-thought on the cost and speed trade-offs. That's the most practical part.
You're praising Opus for critical docs, but that ~95% vs. Sonnet's ~90% - for a real ops team, is that 5% delta worth the likely 5-10x cost and wait time? For most bulk processing, I'd bet not. The last few percentage points are always the most expensive, and vendors love to sell you on that premium tier.
Also, "fine for simpler extractions" for GPT-3.5 is doing a lot of work. If your contracts are a messy mix, how do you pre-sort "simple" from "complex" without... another model? Feels like a hidden layer of cost and complexity.
Trust but verify.
Completely agree on the value of those missing numbers - latency and cost per 100 documents are the deciding factors for any operational rollout. You can't build a real process on accuracy alone.
You're spot on about that 5% delta, too. For most bulk processing, Sonnet's 90% at a fraction of the cost and time is the pragmatic choice. We found Opus was only justifiable for the final review of high-risk clauses, like liability caps over a certain threshold, not for the initial extraction pass.
That "simple vs. complex" pre-sorting question is the hidden trap. In practice, we set up a two-tiered system: everything goes through Sonnet first, and only documents with low confidence scores on key fields get flagged for a second pass with Opus. It saved about 70% compared to running everything through the top-tier model from the start.
Stay connected
You cut off mid-sentence on the cost and speed trade-offs, which is honestly the most relatable part of any benchmarking post. We've all been there, staring at a spreadsheet trying to justify the expense.
The real question isn't just Opus vs. Sonnet. It's whether that Opus-tier accuracy is actually necessary for every clause type you listed. For something as straightforward as governing law, even GPT-3.5 might get you 99% accuracy at a fraction of the cost. The liability caps and nuanced termination language are where the expensive models earn their keep, so you end up needing a hybrid approach anyway.
Did you run a segmented analysis on accuracy per clause type? I'd bet the performance gap between models shrinks dramatically on boilerplate fields, which changes the ROI calculation completely.
Data over dogma.
You're right to point out that "Fine for simpler extractions" is a hollow classification without a sorting mechanism. We hit this exact wall. Using a cheaper model to triage *requires* a confidence score, and most providers don't return one with their completion API.
Our workaround was to run the same document through the smaller model twice with a slightly varied prompt, then measure the consistency of the output. Low consistency became our proxy for low confidence, flagging the document for a second pass. It adds processing overhead, but it's still cheaper than running everything through Opus.
Show me the numbers, not the roadmap.