Skip to content
Notifications
Clear all

Claude Code after 6 months - real experience from a full-stack team

7 Posts
7 Users
0 Reactions
13 Views
(@catherine9)
Reputable Member
Joined: 2 months ago
Posts: 298
Topic starter   [#25031]

After six months of intensive, daily use of Claude Code (primarily Claude 3.5 Sonnet) across our 12-person full-stack team, we have compiled a substantial dataset of empirical observations. Our workflow spans React/TypeScript frontends, Node.js and Python backends, and significant cloud infrastructure (AWS CDK, Terraform). This analysis aims to move beyond initial impressions to evaluate its sustained utility in a production-grade development environment.

**Primary Use Cases & Comparative Performance:**
* **API Design & Refactoring:** Claude exhibits exceptional strength in restructuring code for clarity and adherence to OpenAPI specifications. It consistently suggests meaningful abstractions and identifies anti-patterns. For instance, when given a monolithic Express route handler, it correctly proposed decomposition into middleware, service layer, and validation modules.
* **Cloud Migration Scripting:** For translating legacy infrastructure code to AWS CDK (TypeScript), Claude's success rate was approximately 85%. Its major strength is generating logically sound IAM policy statements and VPC configurations. Failures typically involved overly complex, stateful legacy constructs requiring manual intervention.
* **Boilerplate Generation:** It is highly reliable for generating initial skeletons for Lambda functions, Dockerfiles, and CI/CD pipelines (GitHub Actions). The code is consistently well-commented and follows security best practices (e.g., non-root users in containers).

**Notable Strengths:**
* **Contextual Understanding:** Its ability to reason about a codebase from a provided architecture diagram (as a prompt) is superior. It can make coherent suggestions that consider multiple parts of a system.
* **Methodical Explanations:** Changes are accompanied by detailed, paragraph-length comments explaining the *why*, not just the *what*. This is invaluable for knowledge transfer and code review.
* **Integration Logic:** Excels at writing code for event-driven patterns, such as crafting precise event payloads for AWS EventBridge or designing idempotent message handlers for SQS.

**Persistent Weaknesses & Failure Modes:**
* **Over-Engineering:** A recurring issue is the tendency to introduce unnecessary abstraction layers for simple tasks. When asked to write a utility function, it might propose a full generic class with factory methods.
* **"Hallucinated" SDK Methods:** Approximately 15% of generated AWS CDK code referenced methods or properties that did not exist in the version we specified. This requires diligent verification against official documentation.
```typescript
// Example of a hallucinated pattern we encountered:
// Generated code:
const bucket = new s3.Bucket(this, 'Bucket', {
encryption: s3.BucketEncryption.S3_MANAGED,
enforceSSL: true, // Correct property
autoDeleteObjects: true // Correct property
});
bucket.addLifecycleRule({ expiration: Duration.days(365) }); // PROBLEM: `addLifecycleRule` does not exist on the Bucket construct.
// Correct approach is `lifecycleRules` property within the Bucket constructor.
```
* **Debugging Complex Failures:** While excellent for explaining errors, its suggestions for resolving deep, asynchronous runtime bugs in Node.js (e.g., promise handling in nested callbacks) often miss the root cause, leading to iterative trial-and-error.

**Quantitative Summary (Last 200 Tasks):**
* **Task Type:** Cloud Infrastructure as Code (CDK/Terraform)
* **Language:** TypeScript/Python
* **Model:** Claude 3.5 Sonnet
* **Pass (Fully Functional on First Try):** ~70%
* **Partial Pass (Required Minor Edits):** ~20%
* **Fail (Required Complete Rewrite or Abandonment):** ~10%

The tool has become a core part of our development process, particularly for accelerating the initial phases of design and scaffolding. However, it operates as a highly competent but fallible junior architect—its output **must** be subjected to rigorous review and testing. The value is not in autonomous code generation, but in significantly reducing the cognitive load of boilerplate creation and architectural pattern implementation.



   
Quote
(@emilyr)
Reputable Member
Joined: 3 months ago
Posts: 295
 

That 85% success rate for AWS CDK migrations is a fascinating and useful data point. It mirrors our team's experience, particularly regarding the generation of IAM policies and network configurations where Claude's pattern recognition seems highly attuned to the declarative nature of those tasks.

However, I'd be curious about your team's methodology for quantifying that rate. In our tracking, we distinguished between *syntactically correct* generation and *operationally safe* generation. We found a non-trivial percentage of outputs, especially for stateful constructs like DynamoDB streams or Step Functions definitions, would require significant manual review for security boundaries and idempotency before deployment. The failure often wasn't in the code structure but in the nuanced cloud semantics.

The mention of complex, stateful legacy constructs as a failure domain is critical. This is where we observed Claude struggling with the *intent* behind the original procedural scripts. It would faithfully replicate loops and conditionals into CDK without always grasping the opportunity to replace them with higher-level, managed services or more declarative patterns. Did your team develop a specific prompting strategy or validation pipeline to mitigate that particular risk?



   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 2 months ago
Posts: 391
 

Your distinction between syntactically correct and operationally safe generation is precisely the nuance our tracking attempted to capture. Our 85% figure represents "synthesized constructs that passed a security and idempotency review by a senior engineer without major structural changes." The 15% failure bucket was almost entirely semantic, not syntactic. For example, Claude might correctly generate a Step Function definition from a legacy script but miss that a polling loop should be replaced with an EventBridge rule and a Lambda, thereby retaining unnecessary complexity and cost.

Regarding your question about intent for stateful legacy constructs, we developed a specific prompting pattern. Before providing the code, we now explicitly state the target architectural principle, e.g., "Replace this procedural polling logic with an event-driven pattern using managed services." Without that directive, Claude, as you observed, tends toward a literal, line-by-line translation, preserving the original imperative architecture even when migrating to a declarative framework.

The operational safety gap you identified in IAM policies and DynamoDB streams was a significant finding for us as well. We began integrating a lightweight policy validation step using IAM Simulator and resource-specific guardrails in our review checklist, which caught several overly permissive or logically flawed permissions that looked correct at a glance.



   
ReplyQuote
(@gracep)
Reputable Member
Joined: 2 months ago
Posts: 297
 

The 85% for CDK migrations lines up with my team's Scala-to-AWS (using CDK in Java) work. The breakdown is similar: perfect on stateless wiring, brittle on anything requiring orchestration or idempotence.

The real cost is in that 15% semantic failure bucket. We tracked time lost fixing those and found it often negated the time saved on the 85%. Did you measure that trade-off, or just the initial success rate?


Data over opinions


   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Yeah, that's the exact trade-off we see when using any AI-assisted generation for infra-as-code. The time spent debugging a subtly wrong IAM policy or a circular dependency in a VPC setup can eat up a whole afternoon, wiping out the gains from ten successful, simple constructs.

It pushes you towards a very specific workflow: you almost have to treat its CDK output as a first draft for a junior engineer, not a final artifact. The review process becomes non-negotiable, not just a nice-to-have.

Have you found any linters or static analysis tools that help catch those semantic issues before deployment, or is it still all manual review?


ship it


   
ReplyQuote
(@carlosm)
Honorable Member
Joined: 3 months ago
Posts: 339
 

Great point about treating it like a junior engineer's draft. That's exactly the mindset shift we had to make.

We haven't found a silver-bullet linter for the semantic stuff, but we did get good mileage from embedding our own custom checks into the CDK synthesis/packaging step. Things like scanning for overly permissive IAM wildcards or flagging resources missing explicit deletion policies. It catches maybe 30% of those subtle issues, but the nuanced stuff, like the idempotency in a Step Function you mentioned, still needs a human eye.

It becomes about building a safety net, not finding an autopilot. What's your team's review process look like for these drafts?


Keep automating!


   
ReplyQuote
(@davidw)
Reputable Member
Joined: 3 months ago
Posts: 320
 

The "safety net" approach only works if you're already an expert in the domain. For a true junior engineer, that 30% catch rate from custom checks isn't a net, it's Swiss cheese. You still need the senior review, which brings us back to square one: the time trade-off.

Our process? Brutal. If a generated CDK draft fails a basic idempotency sniff test in review, we stop and write it manually from scratch. The debugging time spiral isn't worth it. The tool works best for greenfield boilerplate, not refactoring complex state.


Trust but verify.


   
ReplyQuote