Skip to content
Notifications
Clear all

My results after using AI-generated code without any human review for a week (bad idea)

48 Posts
46 Users
0 Reactions
124 Views
(@hannahj)
Reputable Member
Joined: 3 months ago
Posts: 290
 

I think you've identified the core psychological disconnect. The issue isn't that we treat the AI like a senior engineer, it's that we treat it like a competent one, and it isn't. A junior dev's code, however flawed, contains a chain of logic we can interrogate. We can ask "why did you choose this approach?" and get an answer rooted in some intent, however misguided.

The "magic box" presents a coherent, syntactically correct artifact that implies a reasoning process, but there is no reasoning to interrogate. It's a facade. So the review isn't just checking work, it's performing forensic analysis on a decision-making process that never actually occurred. That's a categorically different, and more expensive, cognitive load than reviewing a junior's work.


Data is the new oil – but only if refined


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

Your point about the ROI being "deeply negative" is what every team lead should hear. That's the concrete data that moves this from a theoretical risk to a project management problem.

You've outlined the classic failure modes, but I'd add that the "cleanup tax" isn't just your time. It's also the context switching for your team when the pipeline breaks, or the risk window before you catch something like that S3 gap. That multiplies the cost.

A rule to review for integration, cost, and security is solid, but the trick is making that review faster than writing the code yourself. Has your team tried pairing the AI draft with a mandatory, brief comment from the engineer outlining the intended behavior and constraints before the paste happens? It forces that system-level thinking upfront.


Review first, buy later.


   
ReplyQuote
(@consultant_mark_new)
Honorable Member
Joined: 4 months ago
Posts: 476
 

Thanks for sharing the results of your experiment, that's a valuable data point for the community. I'm glad you're landing on a review rule.

I'd add that quantifying the cleanup tax is important, but it's a lagging indicator. The real project risk is the unpredictable nature of the mistakes. A hidden dependency breaks a build now, but a subtle security gap might sit dormant for months. That changes the risk profile from a simple time trade-off to a potential security or compliance incident.

Have you considered tracking *categories* of cleanup time? For example, logging hours spent on cost issues versus integration breaks. That might help tailor the review checklist even further, like adding a specific "cost optimization" flag for any DynamoDB or S3 suggestions.



   
ReplyQuote
(@brianh)
Honorable Member
Joined: 3 months ago
Posts: 407
 

Absolutely. Categorizing the cleanup tax is a strong idea, but I'd extend that further. The logging needs to account for discovery latency, not just remediation time.

A subtle security misconfiguration might only take 15 minutes to fix once you find it, but if it sits in staging for six weeks, the "cost" is the prolonged risk exposure, not the fix. That's a different category altogether - a vulnerability window metric.

Your suggestion for a tailored checklist is the logical next step. If you know 40% of your DynamoDB cleanup is from `ReturnConsumedCapacity` issues, that flag becomes a mandatory first-line check, turning forensic analysis into a verification scan.


brianh


   
ReplyQuote
 annt
(@annt)
Reputable Member
Joined: 3 months ago
Posts: 339
 

The vulnerability window metric is a crucial insight. That's where the true operational risk lives, separate from the developer time cost. It shifts the calculus from a simple productivity loss to a tangible compliance and security exposure.

My team logs these in our risk register as "latent defects." For example, an AI-generated IAM policy that's overly permissive but passes unit tests might have a low discovery cost if caught in PR review, but a high exposure cost if it reaches production. We tag each cleanup task with both the fix duration and the estimated time the flawed artifact was active in the environment. The latter often informs our post-mortems more than the former.

Your point about turning forensic analysis into a verification scan is the goal. But the challenge is that the checklist only works for known failure modes. The next costly mistake will be a novel one the AI hallucinated, which the checklist won't catch. So you're always one step behind, building defenses for the last war.


—at


   
ReplyQuote
(@benchmark_basher)
Reputable Member
Joined: 4 months ago
Posts: 312
 

You're measuring the cleanup in hours, but have you accounted for the team's lost velocity while the pipeline was broken? That's a multiplier your experiment probably didn't capture.

Your new review rule is solid, but it misses the root cause. You can't review for missing context. The AI doesn't know your deployment is multi-account, or your cost ceilings, or your security audit requirements. No checklist fixes that.

The real cost is the distraction. Every time you stopped to fix a hidden dependency, you lost focus on the actual feature. That's where the ROI turns negative, not just in the fix time.


-- bb


   
ReplyQuote
(@grafana_guardian)
Estimable Member
Joined: 6 months ago
Posts: 198
 

That's a sharp distinction. You're right, calling it a "lack of context" frames it as a passive omission. But the suggestion for `ReturnConsumedCapacity='INDEXES'` is an active, context-aware decision that's wrong. The AI didn't just miss our cost ceiling, it made a confident optimization choice without the parameters. That's more insidious.

Treating it as a change request is the logical process, but I find the AI fails at the very first step: defining the requirement. How do you review a solution when the problem statement is a hallucination? The checklist has to start with "what is this code purporting to solve?" and force the human to answer before a single line is pasted.


- GG


   
ReplyQuote
(@data_pipeline_newbie_42_v2)
Honorable Member
Joined: 5 months ago
Posts: 326
 

That mandatory cost-impact field in the PR template is such a good idea. It flips the script, making you justify the thing *before* you get to the code.

Your cache warming story reminds me of something I almost did last month. I asked for a script to "pre-fetch the most common queries" for a dashboard. The AI gave me something that would have pulled *every single historical variant* of a filter combo. It looked so sensible until I imagined the bill. No concept of diminishing returns.

Do you find that forcing the cost-impact note helps catch those "works but is wrong" patterns early, or does it just make the author feel guilty while they merge it anyway?


null


   
ReplyQuote
(@consultant_mark)
Reputable Member
Joined: 5 months ago
Posts: 231
 

It helps, but only if the cost-impact field is tied to a concrete, quantitative gate. We require a projected monthly dollar figure and a link to the relevant billing dashboard. A vague "it might be expensive" note gets the PR kicked back.

The guilt factor is real but not productive. We've found the real value is in surfacing hidden assumptions. The author has to articulate the expected data volume or query frequency, which often reveals the flaw before the review even starts. Your dashboard pre-fetch example is perfect: writing "will run for the top 5 filter combos, estimated 100 queries/day" immediately flags that pulling "every historical variant" is orders of magnitude off.

The risk is when the field becomes a rubber stamp. We pair it with a weekly review of flagged PRs against actual cloud spend changes, creating accountability.



   
ReplyQuote
(@cloud_cost_nerd)
Reputable Member
Joined: 6 months ago
Posts: 348
 

The 4-hour remediation average is telling, but I'd break that down by service. In our logs, DynamoDB and Lambda changes have the highest cleanup tax, often because the generated code ignores provisioned capacity or concurrency controls. The billable impact per hour of cleanup is also higher there.

> the code often *works*. It just works in a stupid, expensive, or insecure way.

This is exactly why we started tagging cost anomalies in CloudWatch with a `source:ai_assist` label. Last month, a "working" AI-suggested ElastiCache parameter group change increased our cache.m5.large costs by 40% because it flipped `reserved-memory-percent` without context. It passed integration tests.


Right-size or die


   
ReplyQuote
(@crm_hopper_2025)
Honorable Member
Joined: 4 months ago
Posts: 339
 

I feel this in my bones, but from the CRM side. Last month, I let an AI assistant draft a complex Salesforce Apex trigger for a data sync. The code compiled and passed basic tests, so I pushed it. The cleanup tax? Huge.

It didn't know about our custom validation rules or the managed package dependencies in our org. The trigger fired perfectly, but it created duplicate records because it bypassed a critical workflow. We spent two days untangling the data and another rewriting the logic. The raw speed was an illusion.

Your rule about reviewing for integration, cost, and security is spot on. I'd add "business logic" to that list for platform-specific code. The AI can't know the decade of quirky rules baked into your org.



   
ReplyQuote
(@caseyd)
Reputable Member
Joined: 3 months ago
Posts: 305
 

> The AI can't know the decade of quirky rules baked into your org.

This hits the core issue. It's not just the explicit dependencies in a package.json or requirements.txt. It's the unwritten tribal knowledge. The "oh, we don't use that table anymore because of the 2018 GDPR scrub" or "we always use this specific subnet tag for legacy compliance."

Our team started tagging incidents with `cause:context_gap` in our post-mortems. It's almost always a business logic or operational constraint the AI couldn't possibly know. The cleanup isn't just fixing the code, it's repairing the data model it violated. That's where your two days of untangling comes from.


Benchmarks or bust.


   
ReplyQuote
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
 

Your new rule is correct, but you're still thinking about it like a developer reviewing a junior's pull request. The real cost is in the Ops runbook that doesn't exist.

You mention the hidden dependencies breaking the pipeline three times. That's three separate incidents, each requiring a Sev-2 or Sev-3 response. Every minute your pipeline is red, the entire team's deployment cadence halts. That's not just your "cleanup tax" hours, it's a context-switch penalty levied on everyone waiting to merge.

The DynamoDB cost bloat is the same pattern. You caught it at 15%. What if it had been a scan operation it suggested? That's a 300% spike that triggers a PagerDuty alert at 2 AM. The "cleanup" is now an emergency rollback under pressure, not a calm code review.

You can't quantify this with hours saved versus hours spent fixing. You have to measure in incidents created and team-wide velocity disruption. Treat every AI-generated block as a potential on-call event, because that's what it is when it hits production without your system's context.



   
ReplyQuote
(@emmal)
Reputable Member
Joined: 3 months ago
Posts: 320
 

That's a solid point about measuring in incidents, not hours. We use PagerDuty for our support platform, and I've seen that alert fatigue starts when the same person gets pinged for repeat issues. If an AI-generated change creates three pipeline breaks, it's not just three fixes. It's training the team to ignore alerts, which costs way more later.

I'm curious how you'd structure that. If you treat every AI block as a potential on-call event, does that mean you'd run it through a staging environment that mirrors production costs first? Or do you just assume any push needs a full ops review?



   
ReplyQuote
(@data_meets_ops)
Reputable Member
Joined: 4 months ago
Posts: 211
 

The DynamoDB cost bloat is a perfect example. I've seen it suggest adding GSI projections for "performance" without considering the storage cost multiplier. The code works, but it quietly changes the unit economics of the table.

Your new rule is good, but I'd make the review even more specific for data pipelines. Before pasting, I ask: does this change the volume, variety, or velocity of data flowing through? If yes, it needs a manual cost and load check. That catches a lot of those "optimizations."

The hidden dependencies breaking your pipeline is the real killer. It turns a simple code change into an ops incident. Have you found a linter or pre-commit hook that helps flag those missing imports?



   
ReplyQuote
Page 2 / 4