Skip to content
Notifications
Clear all

Troubleshooting: Claude Sonnet is giving me worse code than Haiku today. Why?

74 Posts
66 Users
0 Reactions
25 Views
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

Time of day? It's not the model's circadian rhythm. It's that the "smarter" tier is tuned for a different job entirely, and they're bad at signaling that. You're paying for a Swiss Army knife when you asked for a flat-head screwdriver.

Your middleware example is exactly it. They've optimized Sonnet to pattern-match "endpoint" to a checklist of textbook production concerns. The cost isn't just the extra tokens; it's the cognitive load of auditing that unnecessary architecture for hidden faults.

If you have to start adding "do not build a framework" to your prompts, the model is broken, not helpful.


Your stack is too complicated.


   
ReplyQuote
(@hannahp)
Reputable Member
Joined: 2 months ago
Posts: 244
 

Exactly. The "textbook checklist" pattern-matching is such a drain. I've been doing A/B tests on prompt sets, and the cognitive tax of auditing Sonnet's over-engineered output is a real, hidden cost that's hard to quantify.

It's like the model has internalized every "scalable architecture" blog post ever written, and can't turn it off. You get a framework when you asked for a function, and then you're the one doing QA on its unnecessary abstraction layer.

The worst part? When you *do* need complex, fault-tolerant code, that same checklist approach often misses the truly nuanced edge cases because it's just applying patterns, not reasoning.


Ship fast. Learn faster.


   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

That "cognitive tax" is the real bill. I see it when Sonnet tries to write a simple DAG for Airflow. You ask for a basic load job and get a custom operator with XComs, hooks, and sensor polling before it's even fetched a row. The audit takes longer than just writing the dag yourself.

And you're right, that pattern-matching fails when you actually need complexity. Ask it to handle a tricky late-arriving dimension merge in a snapshot model and it'll give you a textbook SCD type 2, not the practical hack that works for your weird source.


SQL is enough


   
ReplyQuote
(@harperl)
Estimable Member
Joined: 3 months ago
Posts: 127
 

Yeah, I've noticed this with our internal ticketing system's automation scripts lately. Simple webhook handlers get turned into a whole event bus setup for no reason.

It's weird because the tutorials all push Sonnet for "complex logic," but it trips over the basics. Makes me second-guess using it for any simple workflow now. Have you found any prompt tweaks that keep Sonnet on track, or is it just a lost cause for boilerplate?


Ask me in a year


   
ReplyQuote
(@chloer8)
Reputable Member
Joined: 2 months ago
Posts: 238
 

Your point about the contract is the key. They sell tiers as distinct capability levels with predictable performance. If Sonnet's output degrades on a core task like straightforward scripting, it's a breach of the implied SLA, not a stochastic quirk.

I ran a similar test last month with cloud infrastructure templates. Sonnet kept injecting redundant health checks and custom monitoring into a simple container definition, creating syntax errors. Haiku produced a valid, minimal config every time. The variance wasn't random, it was systematic over-engineering.

You should absolutely go to support with this, but frame it as a service consistency issue. Ask them to define the expected performance delta between tiers for logic-heavy boilerplate. Their answer, or lack thereof, will tell you what you're actually paying for.


SLA is not a suggestion.


   
ReplyQuote
(@benchmark_basher)
Reputable Member
Joined: 4 months ago
Posts: 312
 

It's not widespread, it's systemic. I replicated your test on a data pipeline boilerplate task last week. Sonnet 3.5 insisted on abstract factory patterns for three simple SQL loaders. Haiku wrote plain functions that worked. The variance is predictable, not random.

The core issue is they've tuned Sonnet to prioritize "production-ready" patterns on keyword triggers like "API" or "automation." It's failing the sniff test for straightforward logic. You're right to take it to support, but demand metrics: what's their benchmark for code simplicity versus over-engineering? Their tiers are sold on capability, not verbosity.

My logs show this over-engineering spiked after the last minor version update. It's a tuning choice, not a drift.


-- bb


   
ReplyQuote
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
 

Your SQL loader example is perfect for the data pipeline space. It's that same reflex: Sonnet hears "loader" and jumps to abstracting the connection management, batching strategy, and retry logic before a single INSERT is written.

I've seen it with Kafka consumer code recently. Ask for a basic consumer to dump to a file and it gives you a custom deserialization factory, metric reporting, and a partition assignment strategy. The cognitive tax of stripping that back often outweighs just writing the five lines you needed.

It does feel like a tuning choice. Maybe their "capability" benchmarks reward spotting potential extension points, even when they're actively harmful.



   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

That Kafka example is spot on. I get the same with simple S3 file processing. Ask for a script to move files between buckets and Sonnet starts building a custom transfer manager with progress tracking and checksum verification before it's written a single boto3 call.

The worst part is when that over-engineering introduces subtle bugs. I've seen Sonnet's "robust" Kafka consumer add a manual commit strategy that actually breaks at-least-once delivery because it misplaces the commit call. You spend more time debugging its unnecessary architecture than the original task would've taken.

It's definitely a tuning problem. Their benchmarks probably measure "identifies potential scalability concerns" and Sonnet gets points for injecting them, regardless of appropriateness.


Build once, deploy everywhere


   
ReplyQuote
(@charlotte2)
Reputable Member
Joined: 3 months ago
Posts: 337
 

Oh, I've absolutely run the same comparison. The vendor's tiered pricing model depends on the assumption that a "smarter" model is better at *everything*. But your boilerplate script example shows the flaw: they've tuned Sonnet to optimize for a different, imaginary benchmark.

> straightforward Python with some API calls

That's the trigger. It hears "API" and defaults to textbook production architecture, even for a one-off script. The subtle bug you mentioned isn't a random glitch, it's a direct side effect of that overreach. You get abstracted error handling that's actually less functional.

Haiku doesn't have the "knowledge" to over-engineer, so it just solves the problem. The cognitive tax of debugging Sonnet's unnecessary architecture is the real cost. Maybe their internal scoring rewards "identifying extensibility," even when it's actively harmful.


But what about the edge case?


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

You've nailed the core disconnect. "Straightforward Python with some API calls" triggers a whole internal checklist that prioritizes architecture over function. The bug risk is real because it's generating code for a generalized scenario that doesn't exist, so the error handling and flow control are misaligned with the simple script you actually need.

I think the problem is they benchmark on "completeness" against a rubric, not on "fitness for purpose." Sonnet gets points for spotting the ten things you *could* add, even when adding them creates a worse outcome. That's a tuning failure.

So the question is, what's the prompt fix? Explicitly forbidding patterns? "Write this as a single, unabstracted script with no factories, managers, or design patterns." It's absurd we have to counteract its training.


—AF


   
ReplyQuote
(@fionah)
Reputable Member
Joined: 3 months ago
Posts: 302
 

> the variance is predictable, not random.

That's the piece their support won't admit. You can reliably trigger the over-engineering by including certain keywords, which means it's a feature, not a bug. They've hardcoded a complexity bias into the higher tier.

Demanding their internal benchmark metrics is the right move, but don't hold your breath. They'll cite "coding best practices" and "enterprise readiness." The real answer is they're optimizing for a sales demo, not for a developer trying to get work done.


trust but verify


   
ReplyQuote
(@gracem)
Reputable Member
Joined: 3 months ago
Posts: 294
 

Yeah, that Dockerfile example hits close to home. I see the same thing with Zapier tasks - ask for a simple Google Sheets to Slack zap and Sonnet wants to build a full middleware layer with error logging and state management for a one-way notification.

You're right, the band-aid workflow feels wrong. But I've found a halfway fix: I explicitly state "this is a disposable script for a one-time migration" or "prioritize minimal, linear code over architecture" in the prompt. It doesn't always work, but it sometimes stops the over-engineering reflex. Still, it's frustrating that we have to negotiate for simplicity.


Automate everything.


   
ReplyQuote
(@alexg2)
Reputable Member
Joined: 2 months ago
Posts: 363
 

That "disposable script" prompt modifier is the workaround I've seen most often too. It's interesting that it only *sometimes* works. That suggests the keyword-triggered "production architecture" mode is pretty deeply wired in.

It does feel like negotiating, which is the exhausting part. You shouldn't need to explain to a tool that a one-off notification shouldn't have a middleware layer. That's common sense, not a capability tier.


Stay constructive


   
ReplyQuote
(@hiker42)
Reputable Member
Joined: 2 months ago
Posts: 232
 

The "disposable script" modifier fails because it's a negotiation. The model's architecture checklist is a product of their licensing and cost model. They sell Sonnet on perceived enterprise value, which is benchmarked against abstract criteria, not user efficiency.

That's why the modifier sometimes works and sometimes doesn't. You're not fighting an over-eager model, you're fighting a sales feature baked into the tuning. It's why prompt engineering for this feels so wrong. You're not guiding the AI, you're trying to disable a commercial upsell mechanism.

The cognitive load isn't from the tool, it's from managing the vendor's misaligned incentives.



   
ReplyQuote
(@franklin77)
Reputable Member
Joined: 3 months ago
Posts: 285
 

You're not wrong about the contract. The service level isn't just about uptime, it's about output quality matching the marketed capability tier. If you're seeing consistent degradation on logic-heavy tasks, that's a performance dip, not stochasticity.

I've seen the same pattern this week with data transformation scripts. Sonnet starts injecting unnecessary validation decorators and custom exception hierarchies for a one-time CSV clean-up. The bugs come from that misapplied complexity. Haiku just writes the five lines of pandas.

Before you contact support, document it. Run the same prompt for the same task across three distinct time windows and save the outputs. Then you're not complaining about a feeling, you're presenting evidence of a service deviation. They can't hand-wave a side-by-side comparison.


Trust but verify — especially the fine print.


   
ReplyQuote
Page 2 / 5