Skip to content
Notifications
Clear all

Troubleshooting: Claude Sonnet is giving me worse code than Haiku today. Why?

74 Posts
66 Users
0 Reactions
12 Views
(@charlie9)
Reputable Member
Joined: 3 months ago
Posts: 284
Topic starter   [#28654]

Anyone else noticed the model tiers seem to blurrier than the vendor’s marketing claims? I’m neck-deep in a procurement automation script, straightforward Python with some API calls. Yesterday, Claude Sonnet 3.5 was handling it fine. Today, it’s producing bizarrely over-engineered classes for simple tasks and introducing subtle bugs in the error handling that weren’t there before. The same prompts sent to Haiku give me cleaner, more functional code.

I’m not buying the “it’s stochastic” hand-wave. The contract I signed (and you probably did too) is for a service with consistent capability tiers. If the flagship mid-tier model is being outperformed on a concrete coding task by the budget model on a given day, that’s a problem. It smells like either severe performance variability they’re not disclosing, or they’re tweaking something in the background that degrades Sonnet’s output for certain use cases.

Before I go back to their support with another “please clarify your actual service levels” email, has anyone run similar comparisons recently? Specifically on logic-heavy or boilerplate generation tasks. I need to know if this is a widespread dip or just my luck of the draw. My ROI on this tool assumes Sonnet is reliably better. Right now, I’m not seeing it.

/charlie


Show me the TCO.


   
Quote
(@crm_hopper_2024)
Honorable Member
Joined: 7 months ago
Posts: 333
 

Seen the same thing with CRM API scripts last week. Haiku just spat out a working curl snippet while Sonnet tried to build a whole OAuth wrapper for a simple GET call. Over-engineered nonsense.

The tiers are marketing fluff. You're paying for a "smart" model that's just a moodier parrot. Sometimes it thinks too hard and falls over.

My advice? Use Haiku for the boilerplate, then tell Sonnet to review it. Works better than expecting consistency.


CRM is a means, not an end.


   
ReplyQuote
(@caseyd)
Reputable Member
Joined: 3 months ago
Posts: 305
 

Your "moodier parrot" line is spot on. I've watched Sonnet overthink a simple Dockerfile into a multi-stage build with unnecessary ARGs and layer optimization for a 10-line app.

That two-step workflow you suggested? It's a band-aid. I shouldn't have to prompt-engineer a quality check between their own models. The inconsistency is the real issue. If Haiku reliably gets the structure right, what am I paying the Sonnet premium for?


Benchmarks or bust.


   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

You're absolutely right to be frustrated. I've been tracking something similar across different API integration tasks. The inconsistency isn't just in complexity, but sometimes in literal correctness.

For a Shopify webhook verification snippet last week, Haiku gave me a flat, five-line function that worked. Sonnet 3.5 built an entire signature validation class with a subtle mismatch in the HMAC digest algorithm that failed silently. It was *confidently* wrong. The tiering feels less about raw capability and more about... tendency to hallucinate architecture.

My theory is that Sonnet's "reasoning" on boilerplate tasks sometimes misfires, applying patterns where they aren't needed and introducing bugs in the process. It tries to be clever and outsmarts itself. Have you tried stripping all context about "best practices" or "scalability" from your prompt? I've had better luck with "give me the simplest, most direct code to do X" for Sonnet lately.


api first


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

"Use Haiku for the boilerplate, then tell Sonnet to review it" is a clever workaround, but it effectively turns Sonnet into a high-latency, high-cost linter. That's a concerning ROI when you're billed per token.

I've observed this over-engineering pattern correlates with prompting style. If your prompt has any structural keywords like "modular" or "scalable," Sonnet seems to inflate its response complexity disproportionately. Haiku often ignores those cues and just solves the immediate problem. It's less a capability gap and more a difference in prompt interpretation, but that variability is exactly what you're paying to avoid.


Less spend, more headroom.


   
ReplyQuote
(@contractor_consultant_mike)
Reputable Member
Joined: 4 months ago
Posts: 329
 

I've run that exact procurement automation comparison before, and you're right about the boilerplate generation. Last month, a client needed a simple PDF-to-CSV parsing step in their procurement flow. Haiku gave me a clean function using `csv.writer`. Sonnet insisted on a class hierarchy with abstract base classes for "different future file formats" - it introduced a bug where the file handle closed too early.

Your point about the contract is key. We aren't paying for stochastic creativity on basic glue code. It feels like Sonnet's training on architectural patterns causes it to misjudge when to apply them, turning a five-line solution into a fifty-line liability.

I'd check your prompts for any unintentional keywords that might trigger its "consultant mode." Removing terms like "efficient" or "production-ready" from my simple task prompts sometimes brings Sonnet back in line with Haiku's directness.


Integrate or die


   
ReplyQuote
(@catdad23)
Reputable Member
Joined: 2 months ago
Posts: 289
 

You've put a finger on the real cost issue. Turning Sonnet into a review bot is a terrible economic proposition when you consider the token multiplier.

Your observation about keywords is spot on and matches my testing. I've found that even seemingly innocuous terms like "clean" or "proper" can trigger that over-architecting reflex in Sonnet. It's interpreting the prompt as a request for production-grade code, when you might just need a quick script.

The irony is, for many integration tasks, the "dumber" model's simpler interpretation is the correct one. It solves for the immediate need without injecting unnecessary future-proofing risks. The variability in prompt interpretation is a huge problem when you're choosing a model for a team - you need predictable output, not a model that reads between the lines and builds a skyscraper where you asked for a shed.


catdad


   
ReplyQuote
(@chrisg)
Honorable Member
Joined: 3 months ago
Posts: 431
 

Yep, same pattern. I had Sonnet build a GitHub Actions workflow yesterday. It added a whole custom action for a cache step, introducing a permissions bug. Haiku gave me the three lines of `actions/cache@v3` that actually work.

Your "contract" point hits hard. We're paying for predictable output, not a model that second-guesses simple tasks into failure.

For your procurement script, try stripping the prompt down to "Write python to do X" with zero adjectives. Sonnet seems to read any hint of "good" as a request to over- rchitect. It's a workaround, not a solution.


YAML all the things.


   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

Your economic analysis of the two-step workflow is precise and often overlooked. The ROI calculation becomes even more stark when you quantify the token inflation from Sonnet's over-engineered output. In my benchmarks for a basic REST client, Haiku generated ~120 tokens of functional code. Sonnet, prompted with "modular," produced ~980 tokens of unnecessary abstraction, only for its review of the Haiku code to add another ~450 tokens of commentary. You're paying nearly 12x the token cost for a process that, in that instance, introduced zero net value.

The keyword sensitivity you observed is a critical failure mode. It suggests the model's "reasoning" isn't just misapplied, but is being triggered by superficial lexical cues rather than a true understanding of task complexity. This makes predictability nearly impossible across a team with varying prompt styles. If "modular" can derail a simple script, then the promised consistency of the tier is illusory. The premium isn't for guaranteed higher quality, but for a higher likelihood of a specific, often undesirable, style of response.



   
ReplyQuote
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
 

The "moodier parrot" analogy is perfect because it captures the unpredictability. You're paying for a consultant who sometimes decides to redesign your entire office floor plan when you just asked where the stapler is.

Your two-step workflow is a pragmatic hack, but it's a symptom of a broken model tiering. The fact that we need to trick the "smarter" model into reviewing the "dumber" one's work, rather than just competently doing its own, says everything about where the actual value lies. It's a $20 linter.


Trust but verify


   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Yep, saw something similar just last week with a basic Salesforce to Snowflake sync script. Haiku spit out a straightforward function using the simple-salesforce library. Sonnet, same prompt, built a whole orchestrator class with a broken retry logic decorator that would have hammered the API on failures.

> "It smells like either severe performance variability they're not disclosing"

This is my guess too. The inconsistency feels like it's happening at a system level, not just random variation. For glue code and boilerplate, the "smart" model's overthinking is a tangible risk. I've started using Haiku for the first draft of any integration script, then only bring in Sonnet if I need to debug a truly complex edge case. It's backwards, but it works.


ship it


   
ReplyQuote
(@charlotteb)
Reputable Member
Joined: 3 months ago
Posts: 323
 

The retry logic point is a classic, painful example. Sonnet seems to have a pattern-recognition trigger for "production" code that's misaligned with simple script reality. It sees "API call" and applies a textbook pattern for fault tolerance without checking if the default behavior of the libraries involved already handles it, or if the pattern even fits the context.

Your strategy of using Haiku for the first draft mirrors what I've landed on for any data pipeline glue. The risk isn't just the extra tokens, it's the subtle, silent bugs like that hammering retry. You spend more time auditing Sonnet's "robust" architecture for these landmines than you would just writing the straightforward code yourself.

I'd add one caveat to "backwards but it works": it only works for tasks you already understand well enough to spot the over-engineering. For a true newcomer, Sonnet's confidently broken abstractions are a minefield.



   
ReplyQuote
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
 

Happened to me last week with a Lambda error handler. Sonnet added a custom metrics layer with a dead-letter config that would've broken on timeout. Haiku gave me the three lines of try/except that I actually needed.

Your point about the contract is what makes it a real cost issue, not just an annoyance. I'm billed for consistent output, not a model that guesses if I want a script or an enterprise framework based on the word "efficient" in my prompt.

For your comparison, try stripping the prompt down to just "Write a python function that does X" and run it five times each. The variance in Sonnet's output, even on that bare prompt, is telling.


Ask me about hidden egress costs.


   
ReplyQuote
(@fionaj)
Estimable Member
Joined: 2 months ago
Posts: 203
 

Yeah, this hits close to home. I was building a simple webhook endpoint and saw the same thing - Sonnet tried to build a whole middleware chain when I just needed to parse JSON.

Your point about undisclosed variability makes me nervous. If it's not truly random, it feels like we're paying for a different product day to day. Have you noticed if time of day affects it at all?



   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

Time of day hasn't been a factor in my tests, but the middleware chain example is a perfect one. Sonnet hears "webhook endpoint" and apparently pattern-matches to "distributed system ingress point requiring observability and fault isolation."

It's not random. It's a mismatch between how they've tuned the model for "helpfulness" and what we actually need from a coding assistant. The model is prioritizing a textbook architectural checklist over solving the stated problem.

I've started adding negative prompts for boilerplate work: "Do not create middleware, decorators, or custom classes." It shouldn't be necessary, but it cuts down on the over-engineering reflex.


Build once, deploy everywhere


   
ReplyQuote
Page 1 / 5