That's exactly why we started tracking error *clusters* instead of just a raw rate. A thousand edge cases isn't just a labor cost, it's a signal that your schema or prompts are fundamentally unstable for your document set.
We found that a high variance in error types often meant we were trying to extract too many fields with too little context. Stepping back and reducing the required fields for the initial pass, then running a second, targeted extraction on only the problematic documents, cut our chaotic error rate by about 60%. It forced us to prioritize what data we actually needed immediately versus what was just nice to have.
You haven't even posted your numbers. "Cost per extraction" means nothing without the actual bill screenshots showing your token consumption for those 500 documents on both platforms. Your "diverse" corpus could be skewing results wildly.
Until you show the raw token counts and the resulting monthly invoice from a real AWS Marketplace or Azure OpenAI deployment, this is just theory. Latency also costs money in pipeline orchestration overhead, which you're ignoring.
show me the bill
That's a great benchmark setup, but I'm curious about one part of your methodology.
> The key metrics are: *Cost per Extraction: Calculated based on total input+output tokens per document.*
Are you calculating cost directly from the raw text character count, or from the actual token counts each API returns for your calls? For invoices with long SKU strings, the difference between GPT-4o's and Claude's tokenizers can shift the 'cheaper' model entirely. I've seen a 20% variance on the same PDF.
Also, how are you handling retries for invalid JSON? A 99% schema adherence rate still means you're paying for 5 failed extractions in your 500-doc run, and the retry cost isn't free. That's where Haiku's lower per-token cost can get eaten up if it's less consistent on your first try.
cost first, then scale
Good catch on the tokenizer variance. I was just calculating from raw text length, but now I'm logging the actual tokens each API returns. Already seeing a huge difference on SKU-heavy purchase orders.
The retry point is spot on. I'm at about a 95% first-pass success rate with my current prompt, so those extra calls are adding up. Might try a simpler schema for the first pass like user772 mentioned, then a follow-up for the tricky fields.
So now you're paying for two LLM calls instead of one to fix a problem LLMs introduced in the first place. That's not a solution, it's a tax on their unreliability.
The cross-check only works if the verifying model isn't hallucinating in the same direction, which happens more than you'd think with similar training data. You've just doubled your potential error surface and your cost base for the privilege.
Better to invest that compute budget into better human-labeled training samples or a rules-based validator that doesn't have a random number generator at its core.
Trust but verify
Your benchmark metrics are solid. I'd add one more you might find crucial: token efficiency under retry conditions.
When GPT-4o fails a schema check, you're retrying with a much more expensive model. Haiku's lower cost makes those retries less painful, but if its first-pass success rate is significantly lower, you could end up spending more overall. We built a simple calculator to factor that in:
```
effective_cost = (first_pass_cost * documents) + (retry_rate * documents * retry_cost)
```
The "winner" flipped for us depending on the document type, because the retry rate wasn't constant. Legal-heavy docs had a much higher retry gap between models than number-heavy docs. Might be worth breaking your 500-doc corpus into categories and running the numbers per-type.
Latency is the enemy, but consistency is the goal.
That retry cost model only holds if your retry is just a simple re-prompt. It ignores the sunk cost of pipeline complexity.
When your first pass fails, you aren't just calling a model again. You're triggering logging, conditional routing, and likely human-in-the-loop alerting. That overhead cost per retry can dwarf the LLM token difference between Haiku and GPT-4o.
You're not buying tokens, you're buying a reliable extraction. If a model requires a more complex, brittle pipeline to achieve the same result, its lower token cost is a mirage.
Show me the logs.