Skip to content
Notifications
Clear all

Just made a list of 20 code patterns Windsurf consistently gets wrong.

49 Posts
47 Users
0 Reactions
5 Views
(@data_diver_43)
Reputable Member
Joined: 4 months ago
Posts: 292
 

That config loading pattern is exactly what I ran into last week. I asked it for a Python script to connect to our internal reporting API, and it gave me something that reads an API key from a plain JSON file on every single request. No environment variable check, nothing.

It makes you wonder what "context" even means to the model. A Lambda handler runs in a specific, constrained environment. The concept of cold starts and keeping things outside the handler seems totally missing from its examples. Have you found a good prompt phrasing to steer it away from that file-reading default?



   
ReplyQuote
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
 

Oh that null guard one is huge, it keeps popping up in data pipeline scripts. Just yesterday I saw it generate a pandas `groupby().sum()` on a column with missing values, where the default behavior just drops those rows silently. That's not a runtime error, it's worse - a silent data integrity failure that doesn't show up until your downstream aggregates are off.

The pattern I'd add to your data handling list is **inappropriate default joins**. It loves to suggest a full inner join for merging datasets without considering if you need to preserve all records from one side. It assumes you always want matching keys, which is rarely true in real ETL work.


Automate all the things.


   
ReplyQuote
(@darrenk)
Honorable Member
Joined: 3 months ago
Posts: 392
 

Oof, that EC2 tagging example is painful. It shows the model doesn't grasp the life cycle of dynamic resources. I've seen the same pattern with it generating scripts that assume database connection strings are static, ignoring that a failover can change the host.

The schema validation gap you mentioned is huge for any external data source. It's not just about types, either. Missing a required field can be just as catastrophic, and a simple JSON schema check would catch it.


dk


   
ReplyQuote
(@diego_h)
Honorable Member
Joined: 6 months ago
Posts: 313
 

I ran into that misapplied idempotency pattern just last week. I was trying to set up a webhook handler to sync user data to our CRM, and it kept generating code that would create duplicate contacts on every retry. It used a simple timestamp check instead of a proper idempotency key with the platform's built-in deduplication.

For someone new to building these integrations, how do you learn to spot these patterns before they cause issues? Is it just experience, or are there specific red flags in the generated code you look for first?


Still learning.


   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

That missing null guard pattern happens constantly with external APIs. You'll get a beautifully crafted map function to process an array of customer objects, but the moment the API returns a partial record with a missing email field, the whole pipeline crashes. It's especially insidious because the test data we use during development is usually pristine.

I've found it consistently picks array methods like `.map` or `.filter` without the defensive `.flatMap` pattern or a simple null-coalescing check. It writes code for the happy path as if it's the only path.

The idempotency issue is the real silent killer in integrations though. Just last month I patched a webhook handler that was creating duplicate Slack channels because it used the channel name as the unique key, not understanding that names can be reused after deletion. The model treated it as a simple "create if not exists" problem without grasping the platform's actual constraints.


api first


   
ReplyQuote
(@clarag)
Reputable Member
Joined: 3 months ago
Posts: 274
 

You mentioned the pagination pattern and it made me think of our project timelines. We tried using a similar script for pulling task data from Asana, and it completely fell over at around 500 records. You're right, it feels like it's built for toy examples.

How do you even start prompting it towards using a cursor? I just end up rewriting it myself every time, which defeats the point a bit.



   
ReplyQuote
(@danielj)
Reputable Member
Joined: 3 months ago
Posts: 254
 

I've had the same struggle with Asana's API! The trick is to explicitly mention the pagination *object* in your prompt. Instead of just "get tasks," try something like "write a script that uses Asana's pagination endpoint, checking for the 'next_page' key in the response object and looping until it's null." You have to force it to think about the response structure, not just the initial call.

It's a bit of a dance, but once you get the phrasing right, it works okay. Still, I end up tweaking the loop condition every single time.


spreadsheet ninja


   
ReplyQuote
(@cloud_cost_owen)
Reputable Member
Joined: 5 months ago
Posts: 181
 

Oh this resonates. I see a version of that data handling pattern all the time in cloud cost automation scripts.

It loves to suggest a `sum()` across line items without filtering out credits or refunds first. The total looks right in a small test dataset, but then you deploy it and your actuals are way off because a single large credit flips the sign. Silent data integrity failure, just like your pandas example.

For the **misapplied idempotency** point - spot on with the UPSERT pattern. I've had to correct so many scripts meant to update Cost Explorer tags, where the generated code would just do a blind `create_tags` call on every run, racking up unnecessary API calls and hitting limits. It never seems to default to checking existing tags first.

Got a list of those 20 patterns? Would love to compare notes, especially on anything related to scheduling or resource provisioning.



   
ReplyQuote
(@fionac)
Reputable Member
Joined: 3 months ago
Posts: 186
 

That "reusable API client with connection pooling" phrasing is exactly what I had to figure out last month. I was pulling campaign metrics and couldn't figure out why the script was so slow until I spotted a new session in a loop.

Have you found that this pattern gap also makes it harder to debug? I spent ages looking at my own logic before realizing the bottleneck was in the client setup it generated.



   
ReplyQuote
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
 

Yes, debugging these performance issues is uniquely time-consuming because you're looking in the wrong place. The bottleneck isn't in your business logic, it's in the boilerplate you trusted it to write.

I've found the same pattern in generated Google Cloud or AWS SDK scripts. It'll instantiate a new client inside a loop processing an item list, completely ignoring the service client's thread-safe, pooled nature. The performance degradation is exponential, not linear, because you're paying the overhead of TLS handshakes and connection setup repeatedly.

You end up profiling your `map` function, when the real fix is hoisting a single `boto3.client('s3')` outside the loop. It's a fundamental misunderstanding of stateful versus stateless client patterns.


benchmark or bust


   
ReplyQuote
(@harryj)
Reputable Member
Joined: 2 months ago
Posts: 381
 

That list is exactly why I treat its output as a rough first draft, not a solution. The aggregation and idempotency gaps you've noted are massive for support automation scripts.

For example, when building a ticket summary from logs, it'll happily write a function to count error types, but completely miss that some log entries might have a null `error_code` field. The script runs fine in dev, then crashes in production.

And the idempotency one hits hard with customer data syncs. I've seen it generate code that creates a new Zendesk user on every sync run instead of checking if the email already exists. You end up with duplicate customer profiles.


Automate the boring stuff.


   
ReplyQuote
(@aidenf)
Reputable Member
Joined: 3 months ago
Posts: 219
 

Exactly - the retry logic without exponential backoff is a huge one on my list. It will dutifully add a `while` loop with a `sleep(1)`, but completely misses the backpressure and rate limit implications. It feels like it learned the concept of "retry" but not the real-world consequences of hammering an API with linear retries.

And you're spot on calling it a **failure to model real-world data entropy**. That's the perfect way to put it. It writes for a clean, ideal dataset every time.

I'll post the full list in a new thread later today - the API integration section is about half of it. The circuit breaker pattern doesn't even seem to be in its vocabulary.


Let the machines do the grunt work


   
ReplyQuote
(@elijahb)
Estimable Member
Joined: 2 months ago
Posts: 201
 

That exponential backoff point hits home. It's not just about being polite to the API, it's about handling the ripple effects in a distributed system. I've seen generated retry logic that, under load, creates cascading failures because every failed request wakes up at the same second to retry, essentially DDoSing the recovering service.

The circuit breaker omission is a glaring blind spot. It's like it understands the concept of "failure" but not "failure state." Without that pattern, you can't gracefully degrade or fail fast when a downstream service is drowning.


Connecting the dots.


   
ReplyQuote
(@emilyj)
Reputable Member
Joined: 3 months ago
Posts: 216
 

The first two patterns you listed are exactly what I've run into with our marketing automation data pipelines. For the aggregation one, I've had it suggest summing up email campaign metrics without checking for undefined values from incomplete data exports.

Do you find these patterns are worse when you're dealing with custom object data in Salesforce, compared to a more standard API? I'm wondering if the training data is skewed.



   
ReplyQuote
(@coffeelover)
Honorable Member
Joined: 3 months ago
Posts: 397
 

The Salesforce API is a mess of custom objects, but that's not the problem. The problem is that all these tools are trained on clean tutorial code, not production garbage data.

It misses undefined values in standard APIs too. Last week I watched it confidently sum a list of CloudWatch metrics where `Datapoints` can be an empty array. It just did `sum(dp['Value'])` and crashed spectacularly.

The training data is skewed toward happy paths, period. It's why the output is brittle.


Just my two cents.


   
ReplyQuote
Page 2 / 4