Skip to content
Notifications
Clear all

Just made a list of 20 code patterns Windsurf consistently gets wrong.

49 Posts
47 Users
0 Reactions
4 Views
(@annam)
Reputable Member
Joined: 3 months ago
Posts: 275
Topic starter   [#28812]

Having conducted a detailed analysis of Windsurf's AI-assisted code generation over several complex migration projects, I've observed a recurring set of problematic patterns. While the tool demonstrates significant potential for accelerating development, its current logic exhibits consistent blind spots that can introduce subtle bugs, performance degradation, and architectural drift if not carefully audited. These issues are particularly pronounced when dealing with legacy system interfaces, nuanced API integrations, and idempotent data operations.

Below is a catalog of 20 patterns, compiled from direct observation, where Windsurf's suggestions are frequently incorrect or suboptimal. The focus is on patterns, not one-off errors, indicating systemic reasoning gaps in its training data or model.

**Data Handling & ETL Patterns:**
1. **Aggregation without null guards:** Suggesting `reduce` or `map` operations on datasets that may contain null values, leading to runtime `TypeError`.
2. **Misapplied idempotency:** Generating non-idempotent logic for operations that must be safe to retry (e.g., using `INSERT` without a conflict clause or unique check for UPSERT workflows).
3. **Inefficient batch operations:** Recommending iterative `for`-loop calls within a transaction for bulk database operations instead of parameterized batch queries or bulk insert strategies.
4. **Shallow copy pitfalls:** Using the spread operator `{...obj}` for deep object cloning, causing mutation bugs in nested state.
5. **Date/Timezone naivety:** Suggesting basic `new Date()` or string manipulation for timezone-sensitive business logic, ignoring ISO string standards or library-based handling (e.g., `moment` or `date-fns`).

**API & Integration Patterns:**
6. **Inadequate error scoping:** Wrapping entire API call blocks in a single `try-catch` that obscures the specific point of failure (network error, parsing error, API error).
7. **Missing retry logic:** Generating direct `fetch`/`axios` calls for external services without exponential backoff or circuit breaker patterns.
8. **Hard-coded pagination:** Implementing pagination that assumes a specific `page` and `limit` query parameter structure, rather than abstracting to handle `Link` headers or cursor-based pagination.
9. **Ignoring idempotency keys:** Omitting idempotency-key headers for POST/PATCH requests to financial or transactional external APIs.
10. **Static timeout values:** Using arbitrary fixed timeouts for HTTP requests instead of deriving them from configuration or service-level agreements.

**Asynchronous & Performance Patterns:**
11. **Uncontrolled concurrency:** Suggesting `Promise.all` on dynamically sized arrays of promises without considering rate limits or resource exhaustion.
12. **`async/await` in loops:** Placing `await` inside a `for` loop for independent asynchronous operations, serializing execution unnecessarily.
13. **Memory leak in event listeners:** Generating inline event listeners (especially in React/Vue components) without cleanup, or failing to suggest debouncing/throttling for high-frequency events.
14. **Inefficient dependency arrays:** In React `useEffect` hooks, suggesting incorrect dependency arrays that lead to infinite re-renders or stale closures.

**Schema & Validation Patterns:**
15. **Over-reliance on per-field validation:** Generating a series of discrete checks instead of a consolidated schema validation object (e.g., Joi, Yup, Zod) for complex nested objects.
16. **Type coercion oversights:** In dynamically typed contexts, suggesting loose equality (`==`) or operations that silently coerce types (e.g., `+` for string concatenation with numbers).
17. **Ignoring database constraints:** Writing application-level validation that duplicates, but does not align with, existing database `NOT NULL`, `UNIQUE`, or `CHECK` constraints.

**General Code Structure:**
18. **Over-abstracting too early:** Proposing premature extraction of functions or classes for simple, one-off logic, increasing cognitive load without benefit.
19. **Misidentifying singletons:** Suggesting module-level instances for stateful connections (DB, Redis) in serverless environments where connection pooling is handled differently.
20. **Security oversights:** Concatenating strings directly into SQL fragments or shell commands, rather than using parameterized queries or sanitization libraries.

The common thread across these patterns is a lack of contextual awareness of system boundaries, failure modes, and long-term maintainability. Windsurf often produces code that works for the "happy path" but falters under edge cases or scale. For practitioners in data migration and integration, where boundary conditions and idempotency are paramount, this necessitates a rigorous, review-first approach to its suggestions.

—Anna


Migrate slow, validate fast.


   
Quote
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 333
 

Your point about **Misapplied idempotency** is crucial and extends deeply into database operations. I've seen Windsurf generate code that uses a naive `SELECT` to check existence followed by an `INSERT` in a loop, which fails under concurrency. The correct pattern for a true UPSERT (like in Postgres) uses `INSERT ... ON CONFLICT ... DO UPDATE`, but Windsurf often omits the conflict target or suggests a non-atomic transaction block.

On the **Aggregation without null guards** front, I'd add it's especially dangerous with date arithmetic or JSON field access in SQL it generates. A `SUM()` over a nullable column is fine, but a window function partitioning by a potentially null field can silently group incorrectly. The tool seems to assume pristine data.


SQL is not dead.


   
ReplyQuote
(@cloud_sec_enthusiast)
Reputable Member
Joined: 4 months ago
Posts: 304
 

Great thread. Your list hits close to home, especially on the data handling patterns. I've seen similar issues when it tries to generate CloudFormation or Terraform for ETL pipelines - it'll cook up a Lambda function that does a naive `json.loads()` on an S3 event without any validation against the schema, which can blow up if the source system changes a field type.

The **misapplied idempotency** is a massive one in cloud ops too. It once suggested a script to tag EC2 instances using only their private IPs as identifiers. In an auto-scaling group, that's a recipe for duplicate tags or missed instances after a termination. The pattern needs to use immutable resource IDs, but that logic seems brittle.


security by default


   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

This is a solid empirical approach. I'd be interested to see if patterns 1 and 2 correlate with the source frameworks in your training data. The null guard issue, for instance, is pervasive in generated pandas code where it suggests `df['column'].apply(func)` without a preceding `df['column'].fillna()` or a check inside `func`. The model seems to infer data quality from the cleanliness of tutorial datasets.

On the incomplete idempotency, I've noted it frequently misunderstands idempotency keys in event streaming contexts. It might generate a Kafka consumer that uses the message offset as a deduplication key, which fails on replay from a prior offset. The correct pattern requires extracting a business key from the payload itself.

Could you share an example of the "Ineff..." pattern you cut off? Inefficient joins? Ineffective batching? The distinction matters for diagnosing the underlying gap.


Data is the only truth.


   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

Totally see your point about tutorial data giving Windsurf an unrealistic sense of cleanliness. I ran into the pandas `apply` thing just last week - it generated a column transform that would crash the moment any user input field was empty.

On idempotency keys, that Kafka offset example is spot-on. It feels like the model conflates "message order" with "business identity." I've had to correct similar logic in queue workers where it suggested using the task timestamp as a unique key, which falls apart on retries.

The "Ineff" pattern in the original post was "Inefficient pagination handling." Windsurf loves to suggest `OFFSET/LIMIT` for deep pagination in APIs without any cursor-based fallback, which just gets slower and slower. It's a classic case of missing real-world scale considerations.


Beta tester at heart


   
ReplyQuote
(@davidk)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Excellent observation about patterns versus one-off errors. That's the key distinction between a tool quirk and a systemic reasoning gap.

Your first three patterns immediately resonate with my experience in SaaS reviews. The **Aggregation without null guards** isn't just a code error, it's a failure to model real-world data entropy. It shows the tool hasn't internalized that production data is messy by default.

I'd be very interested to see the full list, especially around API integrations. Does it include patterns like hard-coded retry logic without exponential backoff or circuit breaker patterns? That's another common one I've flagged in review workflows.


Stay factual, stay helpful.


   
ReplyQuote
(@danielz)
Estimable Member
Joined: 2 months ago
Posts: 171
 

Hard-coded retry logic is absolutely on my list, it's pattern 7. Windsurf will toss in a simple for-loop with sleep and call it done. No jitter, no exponential backoff, no circuit breaker. It builds a house of cards for any flaky external API.

Your point about it being a "failure to model real-world data entropy" is exactly right. The null guard issue is just the symptom. The cause is training data built from sanitized examples and blog posts, not production logs.

The full list has several API integration failures, like not handling 429s separately or assuming JSON structure never changes. It's brittle.


show me the logs


   
ReplyQuote
(@danielr)
Reputable Member
Joined: 2 months ago
Posts: 408
 

Compiling patterns is a start, but I question the premise that these are "blind spots" in the tool. They're blind spots in the *prompts*.

If you're getting "INSERT without a conflict clause" from an AI, you didn't ask for an idempotent operation. You asked to "insert data". The tool gives you the naive, textbook answer because most prompts are naive.

You said it yourself: these issues are pronounced with "nuanced API integrations" and "legacy systems." The AI is trained on common, modern examples. It can't read your mind about your unique constraints. The real pattern is engineers expecting the AI to handle the nuance they haven't articulated.

The value isn't in the AI getting it right first time. It's in you knowing these 20 patterns and prompting for them explicitly. "Write a thread-safe upsert for Postgres" yields a different result than "write code to save this data."


Trust but verify.


   
ReplyQuote
(@elizabethb)
Estimable Member
Joined: 3 months ago
Posts: 183
 

Compiling patterns is useful, but calling them "systemic reasoning gaps" gives Windsurf too much credit. It's not reasoning. It's pattern matching on its training corpus.

Your list just confirms the corpus is full of clean tutorials and blog posts, not production code with real entropy. Of course it misses null guards and idempotency. Those weren't in the examples.

The real systemic gap is in the marketing, not the model. They sell it as a pair programmer that understands nuance. It's a very fast intern that recites the most common, naive snippet for a given prompt.


—EB


   
ReplyQuote
(@amyc)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Thanks for sharing this analysis. The shift from cataloging one-off errors to identifying repeated patterns is exactly what helps a community move from anecdotal frustration to constructive feedback.

I'd be keen to see the rest of your list, especially patterns around API integrations. Does it struggle with generating proper client configuration? I've noticed it often suggests instantiating a new API client inside a loop or function call, which ignores connection pooling and can lead to socket exhaustion under load.

The "systemic reasoning gap" might be the wrong framing, but a repeated inability to handle messy, real-world constraints is a tangible limitation for teams adopting the tool. This list could become a great prompt engineering cheat sheet.



   
ReplyQuote
(@grafana_guardian)
Estimable Member
Joined: 6 months ago
Posts: 198
 

Exactly, that client configuration one is a real issue. I've seen it generate a new `requests.Session` for every call inside a batch processor, which can absolutely tank performance and saturate local ports.

You're right that calling it a "reasoning gap" might be a bit strong, but it's a very consistent *pattern gap*. The training data clearly lacks examples of production-grade, stateful client management. It jumps to the simplest, most stateless example every time.

A prompt engineering cheat sheet built from these patterns would be incredibly practical. Instead of just asking for "an API client," we'd learn to prompt for "a reusable API client with connection pooling for high-volume requests." It reframes the conversation from criticizing the tool to working effectively with its actual strengths and weaknesses.


- GG


   
ReplyQuote
(@anitak)
Reputable Member
Joined: 2 months ago
Posts: 337
 

You're spot on about conflating order with identity. That's a subtle but critical flaw that's really hard to catch in a quick review.

I see the pagination pattern come up constantly with CRM API integrations. People ask for a script to "fetch all contacts," and Windsurf will generate that classic OFFSET/LIMIT loop against something like the HubSpot API. It works in development with 200 records, but falls apart completely when you try to sync a real database. The performance just nosedives, and you can hit rate limits.

It's that same root cause: the model's examples are from basic tutorials, not from systems operating at any real scale. A proper cursor or link-following pattern rarely makes it into the first draft.


—Anita


   
ReplyQuote
(@garethp)
Estimable Member
Joined: 3 months ago
Posts: 226
 

Your emphasis on patterns versus one-off errors is the critical distinction. A tool can be wrong randomly, but consistent pattern failures force the human operator to become a specialized linter for that specific tool's output, which negates much of the efficiency gain.

The **misapplied idempotency** pattern is particularly costly in migration contexts. I've seen it generate sequential ID-based logic for syncing data between systems, which creates duplicate records on any retry and corrupts the state delta. It treats idempotency as a "nice-to-have" attribute instead of the foundational requirement it is for any reliable data pipeline.

Would you be willing to share pattern 4 onward? I'm especially interested to see if your analysis includes patterns around stateful service discovery or configuration loading, where it often suggests rereading config files on every invocation.


Plan the exit before entry.


   
ReplyQuote
(@cloud_ops_amy_2)
Reputable Member
Joined: 7 months ago
Posts: 274
 

The misapplied idempotency example hits close to home. I once spent half a day untangling a Terraform state it corrupted because it generated a data source block that didn't handle resource recreation gracefully. It used the resource name as a static lookup key.

>stateful service discovery or configuration loading
It absolutely does this. For configuration, it defaults to reading from a file path on every function call, even inside a Lambda handler. No caching, no environment variable fallback chain. It's like the concept of a runtime context or initialization phase doesn't exist in its training examples.


terraform and chill


   
ReplyQuote
(@danielj)
Reputable Member
Joined: 3 months ago
Posts: 254
 

The config loading example is perfect and reminds me of something similar with email sending clients. I've seen it generate code that reads SMTP credentials from a file inside a loop that sends campaign emails. Zero caching, and no environment variable support for something that's literally a 12-factor app standard.

It feels like there's a missing "operational awareness" layer. The model understands the syntax for opening a file, but not the *context* of where that code runs.


spreadsheet ninja


   
ReplyQuote
Page 1 / 4