Skip to content
Notifications
Clear all

Just made a list of 20 code patterns Windsurf consistently gets wrong.

49 Posts
47 Users
0 Reactions
6 Views
(@emilyk)
Reputable Member
Joined: 3 months ago
Posts: 286
 

Your initial three patterns highlight the core weakness: it's modeling the **syntax** of a pattern, not its **semantic intent** in a production environment. The null guard omission is especially telling because it's not a complex concept; it's a basic data contract failure. I've quantified this in pg_stat_statements logs, where a single unguarded aggregation can cascade into hundreds of prepared statement errors under load, because the generated code often lacks proper error boundaries for each row processed.

Your second point on **misapplied idempotency** goes beyond API calls. I've observed the same flaw in generated DDL for schema migrations. It will produce an `ALTER TABLE ADD COLUMN` statement without checking `IF NOT EXISTS`, causing the migration to fail on re-run. This indicates the model associates "idempotency" with data operations, not with structural operations, which is a critical architectural oversight.


Show me the numbers, not the roadmap.


   
ReplyQuote
(@data_pipeline_newbie_42)
Reputable Member
Joined: 6 months ago
Posts: 211
 

That timestamp check trap got me too on my first Airbyte pipeline. The red flag for me is any retry logic without a unique key from the source itself - like a message ID or event hash.

I started keeping a checklist next to my monitor:
- Is the dedupe key from the source system or our code?
- Does this API have built-in idempotency keys?
- Are we checking existence *before* create?

For webhooks specifically, I always look for the `X-Request-ID` or similar header now. If the generated code isn't using that, it's probably making the timestamp assumption.

What CRM were you syncing to? Curious if some platforms make this easier than others.



   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

This is so validating, I've been keeping my own mental list! Your first pattern about **aggregation without null guards** hits on a massive issue for RevOps scripts.

I see it constantly when Windsurf generates a forecast rollup from our CRM opportunity data. It'll write a slick-looking function to sum `Amount` fields across a pipeline stage, but completely ignore that a `CloseDate` might be `null` for early-stage deals, or that a custom probability field could be undefined. The script runs fine in the sandbox, then explodes in production the first time it hits a partially imported lead record from a new source.

It feels like it's coding for the perfect, fully-qualified Salesforce record that only exists in Trailhead modules, not the messy reality of a live org with half-completed imports and legacy custom fields.


Pipeline is king.


   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

Totally feel this with the CRM data. Even basic stuff like a `Custom_Probability__c` field being blank throws off a whole forecast sum if the guard isn't there. Does your team just add null checks to every single field manually now, or have you found a better way to handle it?


Still learning


   
ReplyQuote
(@benjislack)
Reputable Member
Joined: 2 months ago
Posts: 244
 

You end up writing a wrapper function that does the null guard for any field path. But then Windsurf doesn't know about your wrapper and keeps generating the raw, unsafe aggregation calls.

It's a band-aid. The real problem is the training data is all clean, perfect objects.


your mileage will vary


   
ReplyQuote
(@danielb)
Reputable Member
Joined: 3 months ago
Posts: 252
 

Look for retry logic that uses timestamps or local counters. That's always wrong for webhooks. A valid idempotency key must be from the source, like `X-Request-ID` or a message payload hash.

Check the target API's docs for native idempotency support (like Salesforce's `Sforce-Auto-Number`). If the generated code doesn't use it, you're in timestamp territory.

It's not about experience, it's about checking two things: source-provided key, platform-native feature. If either is missing from the code, it'll duplicate.



   
ReplyQuote
(@claraj)
Reputable Member
Joined: 2 months ago
Posts: 342
 

The "toy examples" problem is real. It's not just pagination, it's that it only recognizes documented patterns like `next_page` tokens. For a cursor, you have to basically feed it the exact API spec in the prompt, including the parameter name and the response field containing the next cursor.

Even then, I've seen it "invent" a cursor field that doesn't exist, because the pattern feels right. You end up debugging its hallucination of the API instead of writing the actual adapter.


Prove it


   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

Yeah, that "invent a cursor" thing happened to me with the AWS Config API last week. I gave it a prompt for listing resources and it generated code looking for a `NextMarker` field, but the actual response had a `nextToken` key. Had to check the docs anyway to see which one was right.

It ends up taking longer than just writing the pagination loop yourself, doesn't it?



   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

It absolutely takes longer, and the cost implication is where it gets real. That extra debug cycle isn't just your time, it's compute time. Let's say you're generating a script to list all unattached EBS volumes in all regions. The pagination logic runs in a Lambda function. If the generated code has a misnamed cursor key, the loop terminates early on the first page, leaving most resources unlisted. You're now paying for a Lambda run that didn't complete the job, and you have incomplete data that could lead to wasted spend on resources you think are cleaned up but aren't.

This is worse than writing it yourself because it creates a false sense of completeness. You have to validate the output count against what the console shows anyway, which means you're reading the docs you were trying to avoid. I've started treating these pagination helpers as a first draft only, a syntax scaffold that must be verified against the actual API spec line by line.


Every dollar counts.


   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

You're right, that false sense of completeness is the most dangerous part. The cost of the incomplete Lambda run is one thing, but the operational cost of the decision you make with that incomplete data is worse. You think you've cleaned up all unattached volumes and then a month later you get a surprise bill.

It also creates a weird workflow where you're forced to become an expert in the API you were trying to avoid, just to verify the generated code. The time you "saved" on the first draft vanishes in the verification phase, and you're left with more cognitive overhead, not less.



   
ReplyQuote
(@contractor_consultant_mike)
Reputable Member
Joined: 4 months ago
Posts: 329
 

You're spot on with **Misapplied idempotency**. This one cost me a weekend last month. A client had a daily sync job Windsurf "optimized." It was supposed to UPSERT new support tickets into their CRM using a unique external ID from the source system.

The generated code looked logical - it had a check for the external ID's existence. But the pattern was brittle: it did a `SELECT` to check for a match, then an `INSERT` if none was found. The issue? Concurrent runs from a retry could slip through between the check and the insert, creating perfect duplicates. The model's training examples often treat these as sequential steps without race conditions.

The fix wasn't just adding a conflict clause, it was understanding that the source system's timestamp wasn't a reliable idempotency key either. We had to implement a separate deduplication queue, which the initial generated pattern completely missed.


Integrate or die


   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

Yes! Debugging generated code is often harder because you're looking at *your* intent, not the tool's flawed assumptions. I spent way too long tracing a slow ETL job only to find it was creating a new Argo CD session per iteration instead of reusing the client.

A shared client is now a mandatory check in our pull request template. Helps catch it before the code runs.


git push and pray


   
ReplyQuote
(@devops_shift_worker)
Reputable Member
Joined: 4 months ago
Posts: 290
 

Yep, debugging the generated code is a different kind of head-scratcher. You're not debugging your logic, you're debugging a *pattern hallucination*. It looks right, smells right, but the performance hit is hiding in plain sight.

My new rule: any time Windsurf gives me a client object inside a loop or a data-fetching function, I immediately check if it's instantiated there or passed in. Saved me last week on a CloudWatch logs scraper that was re-establishing the session every 10 records. The latency graphs looked like a staircase 😑


NightOps


   
ReplyQuote
(@carlam)
Reputable Member
Joined: 2 months ago
Posts: 234
 

That client-in-loop pattern is brutal. I benchmarked a similar fix last week with the GitHub API client - moving it outside cut average script runtime from ~45 seconds to under 3 for listing repo issues across a big org. The overhead isn't just TLS, it's all the client-side config and auth refreshing it redoes every time.

It makes me wonder if we need a standard checklist for these generated SDK snippets, like "hoist clients, check pagination keys, verify idempotency". What else would you add?


Benchmarking my way to better decisions


   
ReplyQuote
(@chloeh)
Estimable Member
Joined: 3 months ago
Posts: 190
 

Totally agree with your breakdown, especially the idempotency blind spot. That one's bitten us with webhook processors - it'll generate logic that assumes a single delivery, but webhooks can fire twice. The retry logic ends up creating duplicate entries in our system.

We've started tagging our prompts with "idempotent: true" as a hack, but it's hit or miss. It still defaults to a simple existence check, not a proper atomic upsert.

And yeah, debugging these is its own special headache.



   
ReplyQuote
Page 3 / 4