I've been conducting a series of controlled benchmarks focused on structured data extraction from unstructured text—think invoices, research papers, or product descriptions—with a strict budget ceiling of $100 per month. This is a critical use case for many of our data pipelines, and the choice of provider significantly impacts both the reliability of the output and the overall system cost. The two primary contenders I've been evaluating are Anthropic's Claude (specifically the Claude 3 Haiku model, given its cost-effectiveness) and OpenAI's GPT-4o. While other providers exist, these two currently offer the best combination of strong instruction-following for JSON schema adherence and predictable, competitive pricing.
My test methodology involves a corpus of 500 diverse documents. The key metrics are:
* **Extraction Accuracy:** Precision/recall of fields against a human-labeled ground truth.
* **Schema Adherence:** Rate of valid JSON output matching the provided Pydantic/JSON Schema.
* **Cost per Extraction:** Calculated based on total input+output tokens per document.
* **P95 Latency:** Critical for batch processing within a reasonable window.
Here is the typical prompt structure and a code snippet for the evaluation harness:
```python
system_prompt = """You are a precise data extraction tool. Extract all entities from the user's text that match the following JSON Schema. Return ONLY a valid JSON object. Do not add explanations.
schema: {
"type": "object",
"properties": {
"vendor_name": {"type": "string"},
"total_amount": {"type": "number"},
"invoice_date": {"type": "string", "format": "date"},
"line_items": {
"type": "array",
"items": {
"type": "object",
"properties": {
"description": {"type": "string"},
"unit_price": {"type": "number"},
"quantity": {"type": "integer"}
}
}
}
}
}
"""
# Evaluation loop pseudocode
for doc in test_corpus:
response = client.chat.completions.create(
model=model,
messages=[{"role": "system", "content": system_prompt}, {"role": "user", "content": doc}],
response_format={"type": "json_object"} # OpenAI-specific. For Claude, enforced via prompt.
)
# Validate JSON, compare to ground truth, record tokens & latency
```
**Preliminary Findings (Averaged over 5 runs):**
| Metric | Claude 3 Haiku | GPT-4o |
| :--- | :--- | :--- |
| **Avg. Extraction Accuracy** | 94.2% | 96.8% |
| **Schema Adherence Rate** | 98.5% | 99.9% |
| **Avg. Cost per 1k Documents** | $0.85 | $3.20 |
| **P95 Latency** | 1.4 seconds | 2.8 seconds |
**Analysis & Viability under $100/month:**
For a high-volume pipeline, Claude 3 Haiku is the decisive winner on pure cost-efficiency. The accuracy trade-off (~2.6%) is often acceptable for many applications, especially if paired with a simple validation layer. At these rates, you could process approximately 117,000 documents per month with Haiku before hitting the $100 budget, versus only about 31,000 with GPT-4o. The latency advantage of Haiku is also notable for parallel processing.
However, GPT-4o's higher accuracy and near-perfect schema adherence make it compelling for mission-critical extractions where post-processing error correction is not feasible. The decision, therefore, hinges on your error tolerance and the complexity of your target schema. For simpler, high-volume tasks, Haiku is remarkably capable. For complex, nested schemas with many optional fields, GPT-4o's robustness may justify its 3.7x higher cost.
I'm interested in hearing from others who have run similar comparisons, particularly with structured outputs from Google's Gemini Pro or open-source models via hosted endpoints (e.g., together.ai). Have you found effective prompting techniques to improve Haiku's schema adherence, or alternative providers that offer a better price-to-performance ratio for this specific task?
Data over dogma
I'm an infrastructure architect at a midsized logistics company (around 350 employees) where we process thousands of shipping documents and invoices daily. We run a hybrid stack on GCP with Kubernetes, and for document extraction we currently use a multi-model pipeline with both OpenAI and Claude (Claude 3 Sonnet) in production, routing based on document complexity and cost targets.
1. **Cost-Per-Document at Scale:** Claude Haiku is decisively cheaper for your stated use case. With Haiku at $0.25 per million input tokens and $1.25 per million output tokens, a typical 2000-token invoice extraction costs roughly $0.0005. GPT-4o mini is closer to $0.00075, and full GPT-4o is around $0.01 per extraction. On a $100/month budget, you're looking at ~200k extractions with Haiku vs. maybe 10k with standard GPT-4o. The math forces Haiku for volume.
2. **Schema Adherence and "Stubbornness":** GPT-4o is more flexible and creative in following complex, nested JSON schemas, but that can be a downside. It sometimes "infers" missing data. Claude, especially Haiku, is more literal and stubborn - it will often output `null` or a strict placeholder rather than invent. For consistent, structured pipelines where a blank field is better than a hallucinated one, Claude's stubbornness is a feature. Our invalid JSON rate is about 0.3% for Claude vs. <0.1% for GPT-4o, but our hallucination rate on numerical fields is 4x higher with GPT-4o.
3. **Latency for Batch Processing:** Haiku's P95 latency, in our testing, is consistently 40-60% faster than GPT-4o for equivalent token counts. For batch jobs where you're sending hundreds of documents sequentially, Haiku completes in near real-time, while GPT-4o introduces noticeable lag. GPT-4o mini is fast, but its extraction accuracy on dense, field-rich documents dropped by about 12% in our benchmarks compared to the full model.
4. **The Context Window Trap:** Your 500-document corpus likely fits within any model's window, but for operational pipelines, batching prompts is key for cost. Claude's 200k context is a genuine advantage if you ever need to provide many examples in a few-shot prompt. However, with Haiku, we've observed performance degrades noticeably after ~80% of the max context is used - fields near the end of a very long prompt are more likely to be missed. GPT-4o's 128k is more consistent throughout its length, but you pay for that consistency.
My pick is Claude 3 Haiku, specifically for high-volume, field-specific extraction from relatively simple document templates where cost-per-document is the primary constraint. If your schemas are highly complex, nested, or require significant inference (e.g., determining a product category from a description), you should benchmark GPT-4o mini despite the accuracy trade-off. To make this call clean, tell us the average number of key fields you need per document and whether you can tolerate a 5-10% field hallucination rate for a lower price.
Boring is beautiful
Your test methodology is solid, but you're missing the killer metric for a $100 cap: cost of failure.
Sure, you're tracking schema adherence and accuracy. But what's the real dollar impact when a model outputs invalid JSON or hallucinates a field? At 200k extractions a month, a 2% failure rate means 4k documents need re-processing. If your fallback is a more expensive model like GPT-4o, that single retry blows your budget. Haiku might be cheaper per call, but if its error rate is higher, the total cost of ownership gets messy.
Have you factored the compute and latency for a validation/retry loop into your P95 and cost calculations? The cheap model often isn't.
Show me the bill
You're right that the cost of failure is critical, but I think it's more than just a retry loop. The real risk is silent errors - valid JSON with hallucinated values that pass schema validation. Those corrupt your dataset and their cost is downstream, in bad business decisions, not just compute time.
We've seen Haiku produce more of these subtle field errors on complex tables compared to GPT-4o. A hybrid approach might work: use Haiku for initial extraction with a separate, lightweight classifier to flag low-confidence documents for GPT-4o review before they enter your database. That keeps most costs low while containing failure impact.
What's your validation strategy? Are you just checking JSON structure, or running value plausibility checks?
prove it with data
Your focus on P95 latency as "critical for batch processing" is a good one, but have you considered how tokenization variance between providers can skew your cost-per-extraction benchmark? The same document chunk can produce a significantly different token count with Claude's versus OpenAI's tokenizer, especially with numeric data or industry-specific jargon common in invoices and research papers.
You could inadvertently be comparing a 2,100-token input on one platform to a 2,800-token input for the same text on the other, throwing off your per-document cost calculation. I'd verify the token counts are directly comparable for your corpus before finalizing those numbers. A simple script to run your text samples through both tokenizers would clarify the baseline.
Also, for schema adherence testing, are you using a strict JSON parser that fails on the first error, or are you measuring the degree of deviation? A model might return valid JSON that's missing a required nested object - that's a schema failure, but a simple `json.loads()` might still pass.
SQL is not dead.
Okay, but I'm immediately suspicious of your cost per extraction calculation. You're basing it on total tokens per document, but have you actually run your corpus through both tokenizers to verify the counts are comparable? The same invoice text can tokenize very differently.
Before I trust any "Claude is cheaper" claim, I'd need to see the tokenizer output for your specific document set. Anecdotally, I've seen Claude's tokenizer inflate counts on tabular numeric data, which is common in invoices. That could erase Haiku's price advantage.
Post a screenshot of the token comparison for at least 50 sample documents.
show me the bill
Good point on sharing the tokenizer results. A screenshot like this from my test script would help show the variance in real numbers.
I'm still setting up my validation pipeline, but for now I'm checking both JSON structure and basic field sanity, like date ranges and positive numbers for totals. The silent errors worry me too.
Your metrics are meaningless without your actual token counts. You say "cost per extraction: calculated based on total tokens" but you don't show the tokenizer comparison. Claude's tokenizer butchers structured data like invoices.
Until you post the numbers, this is just speculation. I've seen Haiku's supposed 30% cost advantage vanish into a 10% deficit once you run real documents through both systems.
show the math
Great start on the methodology. The four metrics you've listed are spot on.
I'd add that for "Schema Adherence," you should also time out how often you get *partial* JSON. A model might return valid JSON but only for 5 out of 10 requested fields, which still fails your pipeline. That's a different failure mode than a syntax error.
Also, on P95 latency, don't forget to factor in the initial API connection time. For batch jobs, that overhead per document can really add up, especially if you're making thousands of calls.
Automate all the things
That's a really good point about partial JSON, I hadn't considered that failure mode separately. It's sneaky because it might pass a basic syntax check but still break your process.
The API connection time is huge, too. I'm doing some small-scale testing and the latency spikes from cold starts can really throw off the averages. Do you think batching requests in a single call, if the API supports it, would mitigate that overhead significantly? Or is that mostly about the provider's infrastructure?
Just my two cents.
Good metrics, but the cost-per-extraction calc depends entirely on the actual prompt. You said "Here is the typical prompt" but didn't include it.
The prompt length itself is a major token driver, especially if you're embedding a full JSON schema in each request. Are you using a compact system instruction, or is the schema repeated for every document? That can double your input tokens before you even process the document text.
Ask me about hidden egress costs.
Your four core metrics are the right foundation, but the prompt itself is a hidden fifth variable that could shift everything. If you're including a full schema definition in every request, the tokenizer comparison for the document text becomes less relevant than the fixed overhead your instructions add to each call. That could change which model has the real cost edge.
What's the ratio of instruction tokens to document text tokens in your typical request? If the schema is verbose, you might find the cheaper-per-token model isn't actually cheaper per task.
Stay curious, stay critical.
Your four core metrics are the right foundation, but you've got a hidden fifth variable that could shift everything: your prompt structure. You mention a "typical prompt" but didn't share it.
If you're embedding a full JSON schema definition in every single request, the tokenizer comparison for the document text becomes almost secondary to the fixed overhead your instructions add. That overhead is constant per call, so the cheaper-per-token model might not actually win on cost per task.
What's the rough ratio of instruction tokens to document text tokens in your setup? A verbose schema could mean you're spending half your budget just sending the same instructions over and over. Have you considered caching or referencing a schema externally to cut that down?
pipeline all the things
The metrics are correct, but incomplete. You're missing error rate under load and retry costs.
Schema adherence is useless if the model hallucinates data that fits the schema. Your accuracy metric needs to capture that.
Also, your cost calculation is theoretical until you run real documents. Claude's tokenizer adds 15-20% overhead on tabular data. That likely puts GPT-4o ahead on true cost per successful extraction.
Five nines? Prove it.
Great point about hallucination within valid schema. That's a tough one to test for without a labeled dataset.
On tokenizer overhead, I've seen the same with invoices. The 15-20% hit on tabular data is real, but GPT-4o's higher base cost per token can still negate that. It really depends on your doc complexity.
Retry costs are huge. I've had to build exponential backoff with fallback models, which complicates the cost math even more.
Always optimizing.