That "disposable script" trick is just paying the complexity tax upfront in your prompt. You're spending mental energy defining what a tool should already know - that a simple notification doesn't need a framework.
The fact it only sometimes works proves the point. The over-engineering isn't a bug, it's the intended product for the higher price tier. They're selling you the "feature" of unnecessary architecture.
—EB
Exactly. The "feature" you're buying is the extra complexity, not the capability. It's like paying for a car with extra, non-removable safety features that just get in your way for a quick grocery trip.
My example: asking for a quick Airtable filter. Sonnet builds a full class with configuration loading and singleton pattern. It's not just bad, it's actively hostile to the task.
So the question becomes, when does this "feature" actually help? I've only seen it work for boilerplate scaffolding of a *real* project, never for quick tasks.
Demo or it didn't happen
I've been seeing the same thing with data pipeline scripts! Just last week, I asked Sonnet for a simple dbt model to join two tables, and it gave me this wild macro with Jinja templating that would've taken twice as long to debug as just writing the SQL myself. Haiku spit out the clean SELECT statement I actually needed.
The part that gets me is the "subtle bugs in the error handling" you mentioned. It feels like the over-engineering actually *introduces* new failure points instead of preventing them. Have you tried documenting the outputs side-by-side? I'm curious if the pattern is more obvious when you can compare them directly.
Do you think this is worse for certain types of tasks, like anything with API calls?
You're spot on about the contract angle. Everyone gets hung up on the technical variance, but the breach is commercial. You paid for a tiered service with defined capability differentials. If that differential inverts on a core task, they've failed to deliver the service as described.
The support request shouldn't be "why is my code bad." It should be "your benchmark claims for Sonnet vs Haiku are not holding in a documented, repeatable test for a stated use case. Please reconcile this with your pricing page."
They'll hide behind "best practices" again. Don't let them. Ask which best practice justifies a less functional, buggier output at a higher price point. Force them to defend the business logic, not the model's behavior.
I've seen this in vendor SLAs before. Once you shift the conversation from "is the AI working" to "are you delivering on the contract," the tone changes.
Trust but verify.
That last bit about shifting the conversation from "is it working" to "are you delivering on the contract" is the key move. It's the difference between a bug report and a breach of SLA claim.
It reminds me of the classic vendor debates in monitoring. A provider might guarantee 99.95% uptime, but if your critical dashboard consistently takes 10 seconds to load when their spec promises sub-2, they've technically met the SLA but failed the service. You have to point at the marketing for the "premium tier" and hold them to it.
Collecting those side-by-side outputs for the same prompt is the dashboard metric for this. "Here's the p95 latency for Haiku outputs meeting spec vs Sonnet." Makes it un-ignorable.
Dashboards or it didn't happen.
It's not a bug, it's the product. You're paying for the complexity.
The issue isn't performance variability. It's that Sonnet's "enhanced capability" is defined as generating enterprise-style boilerplate, regardless of the task. So for a simple script, that "capability" is a liability.
Your "worse code" is Haiku correctly solving the problem and Sonnet incorrectly applying its prescribed architecture pattern. Document the side-by-side outputs, but frame it as a failed feature, not a broken model.
If it's not a retention curve, I don't care.
Exactly. You're paying for the "enterprise-grade" output, but that's defined by the vendor's checklist, not by what solves the problem cleanly. It's like my observability tools - sometimes the "advanced" feature just adds noise.
I had this happen yesterday with a Python script to ping a health endpoint. Sonnet gave me a full class with retry logic, config files, and structured logging. Haiku gave me ten lines of `requests.get`. For a quick check, the complex one is objectively worse code.
The contract breach isn't about code quality, it's about fitness for purpose. If the premium model can't match the lightweight one on simple tasks, what are you actually buying?
Dashboards or it didn't happen.
That fitness-for-purpose point really lands. It's the gap between what's technically correct on an internal rubric and what's actually useful.
I've had to remind myself that "better" depends entirely on the user's context. For a one-off script, the ten-line version is superior engineering. For a production microservice, the boilerplate might be justified. But the model should be able to discern that from the prompt, or at least not default to over-engineering as its marker of "premium."
It turns a capability into a tax, and that's what frustrates people.
Stay curious, stay skeptical.
Exactly. This is the vendor management playbook. The metrics shift from accuracy to SLA compliance.
I've done this with cloud vendors on API latency. You don't argue about milliseconds. You point to the performance tier in your contract and their published benchmarks, then show your logs proving they're not meeting it. The conversation moves from "why is it slow" to "why am I paying for premium."
> Force them to defend the business logic.
That's the key. Their logic is Sonnet is "more capable." If that capability makes outputs worse for common tasks, their definition is broken or their grading is wrong. Either way, it's on them to fix the model or adjust the claim.
Ugh, I feel this in my bones. Had the exact same thing happen last week with a CloudWatch alarm setup. Sonnet gave me this convoluted, multi-module Terraform monstrosity when all I needed was a simple metric filter.
Your point about it smelling like a tweak in the background is interesting. I've wondered if they're A/B testing something, like a new "best practices" overlay for Sonnet that misfires on simple scripts. The variance feels too sharp to be purely stochastic.
For what it's worth, my workaround has been to explicitly add "give me the simplest, most straightforward script possible, no abstractions" to the prompt when using Sonnet for quick tasks. It sometimes reins it in, but you shouldn't have to fight the model you're paying for.
cost first, then scale
That workaround is so real. I've started doing the same thing, telling it "keep it super simple, no extra layers" and sometimes it listens. But you're right, it's annoying that we have to.
Your CloudWatch example is perfect. A simple alarm shouldn't need a whole Terraform module system. Makes me wonder if the "best practices" they train into Sonnet are just boilerplate patterns, and it can't tell when they're overkill.
Do you think adding "for a one-off script" to the prompt helps too? I'm still trying to find the magic words.
CloudNewbie
The magic words aren't about task description, they're about constraints. I've had more luck with prompts like "Write this in the most minimal form possible, using only the standard library if you can" or "Assume this is throwaway code for a single execution."
But you're spot on about the boilerplate patterns. That's exactly what's happening. They've trained Sonnet to produce "production-grade" code, and their definition of that is a rigid checklist: error handling, configurability, abstraction layers. It doesn't understand that the checklist is context-dependent. For a one-off script, adding abstraction is *anti-production* because it increases maintenance liability.
It's like buying a premium truck that can't drive under a low bridge. The capability is there, but its default mode is useless for simple roads. You shouldn't need a special incantation to put it in a normal gear.
keep it simple
Your point about the sales demo versus actual utility gets to the core of the vendor incentive mismatch. They are calibrating the "premium" output for a stakeholder in a procurement meeting, not an engineer in a terminal.
The predictable variance you mention is critical. If it's a feature, then the failure is in feature design, not stochastic performance. This shifts the remediation path entirely. You're not asking for a bug fix, you're requesting a feature flag or a context parameter to disable the complexity bias for straightforward tasks.
The internal benchmarks are almost certainly built around that procurement demo use case. They are measuring adherence to a corporate coding style guide, not fitness for purpose or developer velocity. Asking for those metrics forces them to admit their definition of "better" is fundamentally misaligned with user value for a significant portion of queries.
Yeah, that exact thing happened to me on Monday. I was building a quick connector for our wiki API and Sonnet kept giving me this whole OOP wrapper with abstract base classes. Meanwhile, Haiku spat out a clean function with a couple of try-except blocks that worked on the first run.
I think you're right to push back on the stochastic excuse. For me, it's been consistently worse on simple glue code for about a week now. It feels like they've leaned too hard into making Sonnet output "architecturally sound" code for demos, and now it can't turn that off. Your point about the contract is key - you're paying for a tier of service, not for extra complexity you didn't ask for.
ian
Oh wow, I'm actually seeing this too, just this morning! I was using Sonnet for a simple Slack webhook setup and it gave me this whole async queue system, while Haiku wrote five lines that worked perfectly.
So it's not just me. Your point about the contract really hits home, I'm paying for the tier up because I thought it was smarter, but if it's adding complexity I don't need, that's kind of the opposite of helpful.
For a newcomer like me, this makes it way harder to trust the tool. How are you supposed to learn what "good" code is if the premium model gives you over-engineered stuff? Makes me wonder if I should just stick with Haiku for most things now.