Skip to content
Notifications
Clear all

My results after a 2-week sprint: Claw wrote code, but we spent a week fixing it.

13 Posts
13 Users
0 Reactions
2 Views
(@annie82)
Reputable Member
Joined: 2 months ago
Posts: 232
Topic starter   [#28762]

Hi everyone, I've been trying to follow the advice here about integrating AI coding assistants to speed things up. My team is small and we use a basic stack: Next.js, a Postgres DB, and a couple of external APIs. We decided to run a focused experiment with Claw (the new one everyone's talking about) on a small but real feature: a user dashboard widget that pulls data from our internal API and formats it with some simple charts.

The promise was huge, right? We gave it detailed specs and let it generate the component, the API call layer, and even the database query. And it did! It produced a ton of code in minutes. We were honestly amazed at first. 🚀

But then we actually tried to integrate it. That's where the "sprint" turned into a fix-a-thon. The generated code used libraries we don't have licensed, the API calls didn't handle our auth pattern, and the DB query, while syntactically correct, was a performance nightmare on our dataset. It looked perfect at a glance, but it was built for a generic "ideal" project, not ours.

So we spent the next week essentially rewriting it. The logic was there, but the integration wasn't. It felt like we got a detailed sketch that we then had to turn into proper construction plans.

My question for you all is... what does this mean for evaluating these tools? I'm overwhelmed. Is the value just in the initial draft, and we should budget equal time for fixing? Or did we do something wrong in our prompts or setup? I'm used to trialing SaaS tools where the trial gives you the real, working product. This feels different. I'd love to hear how you're measuring the actual time savings, if any, when the integration cost is so high.

✌️ annie



   
Quote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Oh, that "looked perfect at a glance" feeling is so familiar. It's the shiny lure on the hook. I had a similar thing happen with an Ansible playbook it wrote for me. It used all the right modules, the YAML was clean, but it assumed a folder structure and user permissions that just didn't exist in our setup. Took me longer to untangle its "perfect" logic than to write the darn thing from scratch with my own messy, understood-in-five-minutes code.

Your point about it being built for a generic ideal project is spot on. These tools don't have the context of your tech debt, your weird legacy auth flow, or that one table that's just built different. They're amazing for greenfield toy projects, but real integration is a whole other beast.

The real value I've found isn't in letting it generate a whole feature, it's in using it like a super-powered junior dev you're pair programming with. You write the skeleton, you define the pattern, and then you say "okay, now write me the function that does X within these constraints." That way you're still driving the bus, but it's handling the tedious bits. Otherwise you're just debugging in the dark.


it worked on my machine


   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

Absolutely. Your pair programming analogy is the only sustainable model I've found. The moment you treat it as an autonomous agent, you're handing over architectural control to something that has zero operational context.

I ran a controlled test last quarter, timing tasks with three approaches: full generation, pair-mode, and manual. The "debug the generated output" phase consistently erased 70-80% of the initial time savings. The pair-mode, where I provided the exact interface and error handling patterns, was net-positive, but only about 15-20% faster than writing it myself, because I spent that saved time on much more rigorous prompt specification.

The hidden cost is in the assumptions, like your Ansible folder structure. It will always pick the most statistically common path, which is often wrong for any established system. You end up reverse-engineering its "ideal" to map it back to your "messy" reality, which is a cognitively heavier lift than just building from your known constraints.


—Alex


   
ReplyQuote
(@emma88)
Reputable Member
Joined: 2 months ago
Posts: 208
 

Pair programming is the only way I'll use these tools now. You mentioned the super-powered junior dev model. That's the key cost control method.

I've started treating the prompt like a contract. If I need an API call, I specify the exact error handling pattern from our existing codebase, down to the variable names. Saves time rewriting style.

But it adds a new overhead. Now I'm spending more time on prompt engineering than I ever did just typing. The trade-off feels worse on smaller tasks.



   
ReplyQuote
(@bobw)
Reputable Member
Joined: 2 months ago
Posts: 342
 

Oof, that "performance nightmare on our dataset" line hits so hard. I've been burned in a scarily similar way with API integrations. It'll generate a clean, textbook GET request with perfect pagination logic, but completely miss that our vendor's API has this quirky, non-standard rate-limiting header that you *must* respect or you get silently throttled. The code works, until it mysteriously starts timing out in production.

It's that gap between textbook and real-world that eats the week. Your "detailed sketch" analogy is perfect. The value is in the draft, but the cost comes from all the institutional knowledge it lacks - your auth flow, your specific Postgres indices, that one API endpoint that's just... weird.

Have you found it useful at all for generating the initial boilerplate, like the React component skeleton, before you step in to wire up the actual data layer yourself? That's been my compromise.


null


   
ReplyQuote
(@gracec)
Reputable Member
Joined: 3 months ago
Posts: 315
 

It's that initial amazement that makes the later adjustment so jarring. You see all that code and think, "We just saved a week." But the reality is you just moved the work from the front end to the back end, from writing to debugging and integrating.

I think the key mistake in your experiment was letting it generate across all three layers at once. That's where the assumptions compound. I've had better luck using it on isolated, pattern-heavy tasks where our team's internal context matters less.

For example, I might ask it to write the database query separately, and then I immediately review it against our actual schema and index strategy. Then I feed that *corrected* query into the prompt for the API layer. You keep it on a very short leash, correcting its context at each step before it builds the next, potentially flawed, layer on top.

It turns the tool into a code suggestion engine rather than a feature generator. The time savings are smaller, but they're real and don't come with a week-long debt.


The right tool saves a thousand meetings.


   
ReplyQuote
(@benjislack)
Reputable Member
Joined: 2 months ago
Posts: 244
 

The short leash approach just turns the developer into a full-time prompt manager and code reviewer. You're spending your time correcting its flawed assumptions step-by-step instead of just writing the correct query yourself in the first place.

Real savings come from automating grunt work, not from generating architecture. Use it for writing unit test stubs or converting data formats, not for designing your data flow. Once you're feeding it corrected context, you've already done the hard part.


your mileage will vary


   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

You're right that feeding it corrected context means you've already done the mental work. That's the core of the efficiency paradox.

But I think the calculus changes when you consider maintenance. If I can get it to produce a unit test stub that follows our specific mocking library patterns, that's not just saving keystrokes - it's enforcing a consistency that pays off for the next person reading the code. The prompting overhead for that is low, because the pattern is repeatable.

Where I completely agree is on architecture. Letting it design data flow is like letting a new hire decide your AWS region strategy based on a blog post. The cost of review and correction utterly negates any benefit.


Every dollar counts.


   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

Totally feel you on the unit test stub example. That's where it shines for us too, but only because we built a custom GitHub Actions workflow that generates them automatically.

It clones our specific mocking patterns from a golden example repo, so the "corrected context" is baked into the automation. The prompt overhead is zero for the dev, which flips the paradox.

But you're spot on about architecture. Letting it design data flow is like trusting a generated PR template to handle your actual merge strategy. It looks right until it meets your real CI gates.


git push and pray


   
ReplyQuote
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

The golden example repo strategy for automating test stubs is clever, it formalizes the "corrected context" into a repeatable benchmark. That's the model I use for synthetic workload generation, where I feed it a JSON schema of our exact table definitions and cardinalities. Without that, the generated queries are useless for performance testing.

But even with that baked-in context, I've found variance in the quality of the output requires a validation step. I run a quick benchmark on any generated query to confirm it hits the expected execution plan. It's cheap, but it's still overhead. Zero prompting overhead doesn't mean zero verification overhead.

Your PR template analogy is perfect. Generated architecture is a nice demo that fails the integration test.


-- bb42


   
ReplyQuote
(@chrisw2)
Reputable Member
Joined: 2 months ago
Posts: 309
 

Yep, the verification overhead is the silent killer. Even with perfect context, you can't skip the validation step.

We tried something similar for generating Prometheus alert rules from JSON spec files. The generated rules looked syntactically perfect and matched our patterns, but we still had to run them through a test Prometheus instance to validate the expression results. The generated queries would sometimes pick weird label matchers that worked but were inefficient.

That final benchmark/validation pass is non-negotiable, and it often eats up the time you "saved" on the initial generation.


Run it yourself.


   
ReplyQuote
(@alexf)
Reputable Member
Joined: 2 months ago
Posts: 233
 

Been there. The initial amazement is the trap.

You let it architect the whole flow across layers. That's where the assumptions multiply. You can't debug three layers of generated code at once.

Treat it like a junior dev who only gets one task at a time. Give it the corrected query, then ask for the API wrapper. Never let it touch the whole stack in one go. The savings vanish when you're untangling its worldview.


Optimize or die.


   
ReplyQuote
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

> The promise was huge, right?

And therein lies the trap. Your sprint result is the canonical case study I see with new teams adopting these tools. That initial amazement at the volume of code is actually the danger signal. It creates a psychological debt, making you feel obligated to salvage the "90% complete" output when a ground-up rewrite is often cheaper.

You identified the core issue: it's built for a generic ideal project. It lacks the institutional memory that defines your actual architecture - your specific Postgres extensions, your API client's retry configuration, even your team's ESLint rules. The time sink isn't just fixing logic bugs, it's reverse-engineering the AI's clean-room assumptions and mapping them onto your messy, real-world constraints. I've found the integration tax is consistently 40-60% of the original development time estimate, which nullifies the headline speed gain.

The pivot is to use it as a pattern-literate codex, not an architect. Feed it a snippet of your existing, production-ready API client and ask it to generate the new endpoint call following that exact pattern. The output then conforms to your world, not its own. This shifts the value from raw generation to templated consistency, which is where the long-term maintenance payoff actually lives.


Mike


   
ReplyQuote