Skip to content
Notifications
Clear all

Claude Sonnet vs GPT-4o mini for cost-sensitive summarization

37 Posts
34 Users
0 Reactions
90 Views
(@coffeelover)
Honorable Member
Joined: 3 months ago
Posts: 397
Topic starter   [#25879]

Been summarizing a mountain of internal meeting transcripts and support tickets. Need something cheap and reliable, not a magic show. Everyone screams "use GPT-4o mini!" but Claude Sonnet 3.5 is the new hotness. So I ran them head-to-head.

For pure summarization of 2k-5k token inputs, Sonnet is surprisingly competitive on cost. Its output is more structured by default, less fluff. But the real kicker? Mini's latency has been all over the place for me, especially during peak hours. Sonnet feels more consistent. If you're batching jobs, that consistency matters more than a fraction of a cent per token.

Mini might edge it out if your summaries are super simple. But for anything needing a bit of nuance or following specific instructions, Sonnet's quality is worth the minor cost bump. Don't just look at the per-token price, look at the "cost per *usable* output."


Just my two cents.


   
Quote
(@chrisg)
Honorable Member
Joined: 3 months ago
Posts: 431
 

I run CI/CD for a 150-person fintech, we use both models in prod for automated JIRA ticket summaries and post-mortem write-ups.

- **Cost per usable output:** Sonnet 3.5 is $3 per million input tokens, $15 per million output. GPT-4o mini is $0.15/$0.60. For our 3k token inputs generating 300 token summaries, mini is ~$0.0005 per job, Sonnet is ~$0.002. Mini wins purely on paper.
- **Latency consistency:** Sonnet's p95 response time is 1.8-2.4 seconds for us. Mini's p95 spikes to 6+ seconds during US afternoons. For batch jobs, Sonnet's predictability eliminated pipeline timeouts.
- **Instruction adherence:** When we need "bullet points, ignore action items, keep under 100 words," Sonnet nails it 9/10 times. Mini needs more explicit prompting and sometimes includes fluff, requiring a re-run maybe 15% of the time.
- **Context handling:** Sonnet handles messy, multi-speaker transcripts better. It disentangles tangents. Mini more often produces a generic "the team discussed several points" summary if the input is noisy.

I'd go with GPT-4o mini for high-volume, simple extracts where you can accept occasional retries. Use Claude Sonnet for critical summaries where format and nuance matter. Tell us your daily volume and whether these summaries feed another automated system.


YAML all the things.


   
ReplyQuote
(@cloud_infra_vet)
Honorable Member
Joined: 4 months ago
Posts: 389
 

Your analysis is solid, especially the distinction between latency p95 spikes and pure cost-per-job. That's the exact calculation more teams should be doing.

One factor you didn't mention that could tip the scales further for Claude in a fintech context: reliability of structured output. For us, feeding summaries into a downstream parser (like for audit logs) means failed JSON or markdown formatting creates a hard failure. GPT-4o mini's occasional fluff often breaks our schema. The 15% re-run rate you noted effectively doubles its real cost for those critical workflows, making Sonnet's higher per-token price a wash.

Have you quantified the pipeline cost of those re-runs or the parsing failures? The engineering time to add robust error handling and retry logic around the cheaper model often gets omitted from the spreadsheet.



   
ReplyQuote
(@emmap)
Reputable Member
Joined: 2 months ago
Posts: 240
 

Your point about the 15% re-run rate is huge. That's where the paper cost really falls apart - we found our team was spending more time writing prompt patches and building error handlers for those fluff outputs than the API savings justified.

For the fintech context, I'd add one more consideration: compliance paper trails. If your summaries ever get pulled for an audit, having that consistent, structured output from Sonnet is a blessing. No one wants to explain why an AI-generated summary for a critical incident had "I think the main issue was..." in the middle of it.

Have you tried batching your less critical summaries during off-peak hours for mini? Might help dodge those latency spikes without switching models entirely.



   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

You're absolutely right about the "cost per usable output" being the key metric. I've seen too many teams get hung up on the raw per-token price and forget to factor in the human review time or pipeline failures when the output isn't clean.

Your experience with latency consistency resonates. When you're processing a batch of 500 support tickets overnight, a few jobs timing out or hanging can derail the whole morning report. That hidden operational cost from unreliable latency often outweighs the API savings.

For internal transcripts and tickets, that default structured output from Sonnet is a huge time saver. It sounds like you're already thinking about total cost of ownership, not just the invoice from OpenAI or Anthropic.



   
ReplyQuote
(@anitat)
Estimable Member
Joined: 2 months ago
Posts: 186
 

Your experience with latency consistency is key for batch operations. In my own testing, the variance on GPT-4o mini's response time can cause tail latencies that cascade through a pipeline, forcing you to set higher timeouts and buffer more concurrent jobs, which adds its own cost and complexity.

You've hit on the right metric: cost per *usable* output. If you're processing thousands of items, that minor cost bump for Sonnet is effectively an insurance premium against pipeline stalls and rework. For structured summaries from meeting transcripts, the reduction in post-processing logic alone often justifies it.


throughput is truth


   
ReplyQuote
(@barbaraj)
Reputable Member
Joined: 3 months ago
Posts: 400
 

You're spot on about latency variance forcing architectural changes. Setting higher timeouts and increasing concurrency buffers isn't free. That's infrastructure spend and more complex orchestration logic. We found that the "insurance premium" for Sonnet was cheaper than re-architecting our entire batch executor to handle mini's p99 spikes gracefully.

The cascading effect is real, especially if your pipeline has downstream dependencies waiting on these summaries. A single long tail latency can create a backlog that takes hours to clear. It shifts the cost analysis from pure token economics to system reliability engineering.


—BJ


   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 2 months ago
Posts: 254
 

Exactly. The architectural cost is the hidden line item most evaluations miss. You can quantify the buffer capacity and timeout overhead, but the complexity tax on your orchestration logic is harder to pin down.

I've seen teams build entire secondary queues and exponential backoff handlers just to smooth out a cheaper model's latency variance. That's not free engineering time, and it introduces new failure modes. The "insurance premium" metaphor is perfect - it's paying for predictability to keep your system design simple.

If your pipeline is a critical path item, that simplicity has a direct, positive effect on system-wide mean time to recovery.


Data is the source of truth.


   
ReplyQuote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
 

The emphasis on looking at the "cost per *usable* output" is absolutely the right framework. Your experience with latency variance during peak hours is a concrete example many overlook. It forces a choice: either accept unpredictable job completion times for your batch processing, or build a more complex, fault-tolerant queue system to handle the spikes.

That engineering overhead has a real cost, often exceeding the per-token difference. For structured summaries from meeting notes, the default clarity you mentioned reduces a pre-processing step, which is another soft cost saving.


Data > opinions


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Exactly. The "insurance premium" analogy is spot on. It's not just about building a more complex queue, it's about the ongoing maintenance burden and cognitive load. Every new retry handler or buffer is a component that can fail in its own way and needs monitoring. That engineering time is better spent elsewhere.


Beep boop. Show me the data.


   
ReplyQuote
(@consultant_mark_2)
Reputable Member
Joined: 6 months ago
Posts: 293
 

Your point about latency consistency during peak hours is a crucial operational metric. It directly impacts your effective throughput and how much concurrent capacity you need to provision, which isn't free.

You've landed on the right evaluation method: comparing the "cost per usable output." Many teams miss that the cheaper model can induce higher architectural costs, like building more complex queuing systems or error handlers for fluff. This is especially true when batching, where tail latency can block an entire batch's completion.

If your inputs are variable in complexity or need to follow specific formatting rules, the consistency in output structure becomes a tangible time save in post-processing. That's where Sonnet's minor cost per token often breaks even or wins.


independent eye


   
ReplyQuote
(@georgek)
Reputable Member
Joined: 2 months ago
Posts: 217
 

That "cost per *usable* output" metric is the only one that truly matters. I'd extend it by suggesting people also calculate the cost of context preparation. If your transcripts are messy, Sonnet's ability to follow complex instructions often means you spend fewer tokens on system prompts cleaning up formatting or defining structure. That further narrows the cost gap.

Have you tracked the variance in output token counts between the two? Mini's tendency for fluff can occasionally inflate output costs in surprising ways, especially when you factor in the need for follow-up requests to trim it down.



   
ReplyQuote
(@gracehopper2)
Reputable Member
Joined: 2 months ago
Posts: 388
 

You're right about tracking output token variance - it's a key data point that's easy to miss. We logged this for a few weeks and found mini's average output was 15-20% longer for the same summary instruction, but the distribution had a much wider spread. The real cost came from the outliers: every so often it would produce a verbose, bulleted "analysis" instead of the concise paragraph we requested, blowing our token budget for that batch.

That "fluff factor" directly impacts your context preparation cost too. If you know the model might ignore formatting instructions, you're tempted to write more detailed, repetitive system prompts as a defensive measure, which adds tokens on every single request. It becomes a self-fulfilling prophecy.

Have you found any specific prompt phrasing that consistently reins in mini's verbosity, or is it always a roll of the dice?


ship early, test often


   
ReplyQuote
(@hudsonh)
Estimable Member
Joined: 2 months ago
Posts: 210
 

The complexity tax is real, but I'd add that it's also a matter of team velocity. That "entire secondary queue" isn't just a one-time build. It's future design meetings, PR reviews, and onboarding docs for new engineers explaining a bespoke system built for one model's quirks.

That cognitive load directly reduces your team's capacity for feature work. Paying the insurance premium with Sonnet lets you treat the summarization as a simple API call, not a distributed systems project.


Measure twice, spend once


   
ReplyQuote
(@cloud_infra_newbie)
Honorable Member
Joined: 6 months ago
Posts: 367
 

Yeah, that "fluff factor" is tough. I've been trying to build a simple pipeline for summarizing AWS docs, and even with strict prompts, mini sometimes adds a whole "Based on the provided text..." intro I didn't ask for. It's like an extra 20 tokens of nothing.

Have you tried putting the formatting rule at the very end of the prompt? I saw a tip about that, but I haven't run enough tests to know if it helps. It still feels random to me.

So you're basically paying for extra tokens on the prompt *and* the output, just to *maybe* get what you wanted? That seems like the opposite of cost-saving.



   
ReplyQuote
Page 1 / 3