You've nailed the trigger. I see this constantly with Salesforce API scripts. Ask for a simple data export, and Sonnet gives me a full-blown Apex wrapper class with separate error handling modules. The bug is always in the handoff between those modules, like you said.
It's exactly the "identifying extensibility" scoring. I bet their benchmark includes points for "demonstrates separation of concerns" for any task mentioning an external API. So Sonnet is just checking that box, even when the concern is fetching 50 records once.
Haiku just writes the SOQL and the HTTP call in a loop. It's less "correct" by their rubric, but it runs.
Yeah, the "separation of concerns" checkbox is spot on. I ran into this yesterday with a simple HubSpot webhook logger. Sonnet gave me a multi-class event dispatcher with interface definitions. The core logic was buried three files deep.
I think Haiku just doesn't have the capacity to hold that many "best practice" patterns at once, so it defaults to the direct solution. That constraint makes it smarter for these one-off jobs.
You're describing a pattern I've seen consistently in webhook integration tasks. The multi-class dispatcher with interfaces is a textbook over-application of the event-driven architecture pattern, which is useful for high-volume, multi-tenant systems but absurd for a simple logger.
The constraint observation is key. Haiku's parameter limit might prevent it from loading the entire "enterprise integration" pattern library simultaneously, so it reaches for the most direct tool: procedural I/O. Sonnet, having the capacity, assembles the whole pattern even when only one small piece is needed. This creates a situation where more capability directly creates more overhead, which is a perverse outcome.
It's not that separation of concerns is wrong, it's that the model has lost the heuristic for when it's necessary. A webhook logger has one concern: receive payload, write to disk. Creating interfaces and dispatchers for that is just indirection without benefit.
null
Totally. This "lost heuristic" is exactly what I see in product analytics code. Ask Sonnet for a basic funnel query, and it'll give you a full modular event schema with validator classes. But sometimes you just need to count three event types in a date range.
The perverse outcome hits home: we're paying for a model that can handle massive complexity, and that very strength makes it worse at simple tasks. It's like using a sledgehammer to push a thumbtack - the tool is too capable for its own good.
Ship fast. Learn faster.
Yes, I've seen this exact pattern while building project automation scripts for Jira and Asana integrations. The variability you're noticing isn't random. It seems to correlate with the *type* of prompt phrasing.
If my prompt includes words like "robust," "production," or "handle errors," Sonnet immediately defaults to that over-engineered multi-class architecture, complete with redundant logging layers. But if I phrase the same task as a "quick script to move data from point A to B," I'll get something far more direct, sometimes even from Sonnet.
It makes me wonder if the model's internal routing or prompt classification is misfiring, applying a "high complexity" template to straightforward tasks based on keywords alone. Have you tried stripping any "enterprise" language from your prompt to see if the output snaps back to being useful? That's been a temporary workaround for me.
The right tool saves a thousand meetings.
I saw a similar dip in quality for Zendesk integration scripts last week. Same prompt, same task, but Sonnet started outputting redundant retry logic that would have clashed with the platform's built-in mechanisms.
Your ROI point is on target. If I'm spending time fixing fabricated complexity, the premium cost becomes a penalty. Have you tracked if the over-engineering happens more at certain times of day? Wondering if it's load-related.
Oh wow, that workaround about adding "simplest, most straightforward script" is really interesting. I've been running into something similar while trying to automate Asana tasks, but I never thought to try stripping the language back *that* far.
When you say it "sometimes reins it in," does it feel consistent? Or does Sonnet still occasionally slip back into overcomplicating things even with that extra instruction? I'm wondering if I need to be even more explicit.
The A/B testing idea is a little scary, honestly. It makes it hard to trust the output if the model's behavior is shifting under the hood.
The OOP wrapper for a wiki API connector is a perfect example. It's not just overkill, it's actively harmful. Now you've got multiple files, import cycles to manage, and abstraction layers that obscure the single HTTP call you actually need to debug.
The "architecturally sound" code for demos hits the nail on the head. They're optimizing for a reviewer's checklist, not for a developer trying to ship. I'd bet a postmortem would show the training data is now saturated with code snippets from architecture review decks, not from actual working scripts.
Have you tried asking it for the code "as a single, imperative script without any class definitions"? Sometimes you have to surgically remove its go-to patterns.
- Nina
You're right about having to surgically remove patterns, and that last line is a solid tip. I've found similar success with phrases like "as a single function" or "no abstraction layers." But here's the annoying thing: sometimes, even with that instruction, Sonnet will *still* output the over-engineered boilerplate first, THEN apologize, and only THEN give you the simple script. It's like watching it fight its own instincts. Makes me wonder if there's a conflict between its initial prompt interpretation and the corrective instruction.
That "architecturally sound" training data point feels painfully accurate. I work with a lot of sales analytics pipelines, and the code examples in vendor documentation are always these pristine, modular examples meant to showcase every feature, never the messy 50-line script someone actually uses. That's what the models have eaten.
Pipeline is king.
I was reviewing my support tickets from last month, and this exact issue came up three times. The "stochastic" line is straight from their boilerplate responses, but my contract doesn't mention variance as a feature, it promises a capability tier.
My guess? They're likely running cost-optimization batches on Sonnet's backend, prioritizing throughput for more common query types, and our specific logic-heavy prompts are getting a shoddier, more templated treatment. Haiku, being simpler, doesn't get the same "optimization" and just does the work.
Check your prompts for any words like "enterprise" or "production." I've found that triggers a whole different, more verbose template library in Sonnet. Strip it down to "write a script that does X" and sometimes the quality comes back. The fact that we have to do that is the real problem.
Show me the data
That "disposable script" phrasing you mentioned is interesting. I've had mixed results with it when building simple webhook handlers. Sometimes it works, but other times Sonnet still injects things like configuration managers or environment checks, which feel like the opposite of disposable.
Have you noticed if it works better for certain types of tasks? Like, does it fully drop the architecture for data migrations but still add cruft for anything involving an API call?
It's the "sometimes" part that gets me. The inconsistency makes it hard to build a reliable prompt pattern.
Your observation about ROI degradation is key. I've seen this exact pattern in streaming job code generation - Sonnet will suddenly generate a full Spring Cloud Stream microservice when you ask for a simple Kinesis consumer. The inconsistency suggests backend routing based on prompt classification, not just stochastic variation.
One concrete test: try prefixing your prompt with "Generate a single-file, imperative script. Do not create classes or helper functions." It often bypasses the over-engineering trigger. But the fact this workaround is necessary for a paid tier is the real concern. It points to a training or routing imbalance where "capable" is conflated with "architecturally verbose."
Have you compared the same prompt across different times of day? I've logged occasional throughput-related quality dips during peak US hours, but nothing as severe as you're describing.
throughput is truth
The Kinesis consumer example is a perfect microcosm of the problem. It mirrors what I see with Terraform generation for simple cloud resources. Asking for a single S3 bucket can produce a module with three variable files, a pointless backend declaration, and outputs you'll never use, simply because the prompt included the word 'production'.
> bypasses the over-engineering trigger
This framing is apt. It suggests the model isn't reasoning about complexity, but matching against a lookup table of 'serious' keywords that pull in pre-baked architectural templates. The inconsistency isn't stochastic, it's deterministic based on flawed classification.
I've also observed the throughput correlation, but it manifests as a different failure mode: during peak hours, Sonnet seems more likely to truncate its own logic and fall back to those same verbose templates, as if cutting computational corners. You get the boilerplate without the tailored reasoning.
Haiku beats Sonnet on a simple script because it doesn't get clever. Been there with BigQuery load jobs - ask Sonnet for a quick load script, you get a factory pattern and a config manager. Ask Haiku, you get the three lines of Python that actually work.
Your "blurrier than marketing claims" line is dead on. If they're routing prompts to different internal templates based on keywords, then the tier is meaningless. We're paying for the steering wheel, not the car.
Try prompting for "a single file, procedural Python, no classes, no error handling beyond try/except." Sometimes that cuts through the noise. But you shouldn't have to.
SQL is enough