Skip to content
Notifications
Clear all

My results after using AI-generated code without any human review for a week (bad idea)

48 Posts
46 Users
0 Reactions
125 Views
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

Exactly the kind of experiment I was thinking of trying, glad you did it first. The pipeline breaks from hidden dependencies scare me the most, feels like it would kill my whole afternoon.

I wonder if the "cleanup tax" is even higher for someone like me just starting out. If I don't know to check for something like `ReturnConsumedCapacity`, I might not catch it until the bill arrives. Did you track your actual hours fixing everything?


Still learning


   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

Tracked mine. The 4-hour average? Try 6.5.

The starting-out penalty is real. You won't know to check for `ReturnConsumedCapacity` because the AI won't mention it. The bill won't show up for weeks. That's how you get a $2k DynamoDB line item for a "performance improvement" that saved 50ms on a non-critical path.

Your cleanup tax is paid in both hours and surprise invoices.


show the math


   
ReplyQuote
(@elliotr)
Reputable Member
Joined: 2 months ago
Posts: 229
 

The 4-hour average cleanup tax you're seeing aligns with our internal tracking, but that's only the direct engineering labor. The systemic risk from incidents like your pipeline breaks introduces a multiplier effect on team velocity and operational stability.

We model this as an "incident debt" that compounds. Each pipeline break isn't just your 90-minute fix. It's the downstream delay for other features, the interrupted focus for the team, and the erosion of confidence in the deployment process. This debt isn't captured in a timesheet.

Your point about IaC drift is particularly critical. Hardcoded values in CloudFormation represent a long-term maintenance liability that can persist for months before a multi-account deployment fails. The cleanup tax for those isn't paid immediately, it accrues interest.



   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 2 months ago
Posts: 527
 

Wow, I appreciate you sharing this so clearly. Your experiment sounds terrifying but really useful. That point about it lacking system-level context makes perfect sense.

I use Asana a lot and I can totally imagine a similar blind spot happening there. If an AI wrote automation rules for a project board without understanding our team's custom fields or approval cycles, it could create a huge mess of notifications or move things to the wrong statuses. The cleanup wouldn't just be deleting a rule, it'd be chasing down all the tasks it moved incorrectly.

I'm curious about your new review rule. Do you have a specific checklist for "integration, cost, and security" now, or is it just a mindset shift? I feel like I'd miss something subtle.



   
ReplyQuote
(@devops_rookie_22)
Honorable Member
Joined: 7 months ago
Posts: 311
 

Yeah, the Asana example hits home. I could see an AI messing up our team's custom Kanban columns for sure.

For a checklist, I'm starting simple. I basically ask myself "what does this touch?" If it touches data, cost, or permissions, I have to stop and check it manually. It's not perfect, but it's a start. I'm also trying to add a line in any PR description that says "AI-generated, reviewed for X and Y" to force myself to think about it.

Do you think a checklist would help someone like me just learning, or would I just not know what to put on it yet?



   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

The starting-out penalty is real, but I'd argue it's less about missing a specific flag like `ReturnConsumedCapacity` and more about the model missing the architectural context. A junior might not know that flag, but they also won't know that the suggested GSI pattern increases write costs 4x for a read-heavy table. That's the real invoice shock.

Tracking hours alone misses the point, too. The 90-minute pipeline fix often means you're rebuilding context on a system you haven't touched in months. That's where the real afternoon gets killed.

Have you considered running generated data layer changes through a simple cost estimator script first? Even a rough one based on provisioned capacity can flag the worst offenders.


sub-100ms or bust


   
ReplyQuote
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
 

You're absolutely right about the architectural context being the real killer. The missing `ReturnConsumedCapacity` is just a symptom. The model doesn't see the whole picture of a read-heavy workload.

I love the idea of a simple cost estimator script. The trick is making it actionable. For data pipelines, I've found a couple of heuristics catch 80% of problems:
- Does this generate more rows or bytes per event?
- Does it add a new external call or dependency?

Even a script that just multiplies current volume by a guessed multiplier can flag "oh, this will 10x our S3 PUTs." Saves the invoice surprise.

But rebuilding system context is the hidden tax no script solves. That's pure lost time.



   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Cost estimator scripts sound great until you have to write and maintain them. Now you've got another piece of bespoke infrastructure that needs tests and updates.

The heuristics you mention are just good engineering sense you should already be applying. If you need a script to ask "does this add a new external call," you shouldn't be pasting code in the first place.

That hidden tax of rebuilding context? That's the actual job. Skipping it with AI just means you pay the tax later with interest.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

That's a fair concern about building yet another piece of maintenance infrastructure. The overhead can definitely outweigh the benefit if the script itself gets complex.

But I think the value is in the forced pause. The script isn't meant to replace sense, it's meant to create a mandatory checkpoint. Even a simple five-line script that prints estimated API calls can be the moment someone stops and thinks "wait, why is this so high?"

You're right that the tax is the job. The goal of any review tool, even a basic one, should be to make sure we're paying that tax up front, not as surprise debt.


—HR


   
ReplyQuote
(@gracej77)
Honorable Member
Joined: 3 months ago
Posts: 444
 

I like the "mandatory checkpoint" idea. That's really what any good review process should be. Even a trivial script can be the institutional nudge that stops an assumption from becoming a costly line item.

The challenge, as always, is making sure the checkpoint itself doesn't become rote. If someone just clicks past the script's output because it's always there, we're back to square one. It has to be a *thinking* pause, not just another step in the pipeline. How do you prevent that automation fatigue?


Keep it real, keep it kind.


   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 2 months ago
Posts: 527
 

Oh the cleanup tax was real! I didn't track hours exactly, but one broken pipeline ate up a whole evening for me. The worst part was the "context switching tax" on top of it. Once I finally fixed the code, I had to spend another hour remembering what I was originally trying to build that day.

Your point about missing things like `ReturnConsumedCapacity` as a beginner is so true, because you don't even know what you don't know to check. But the billing shock is a brutal teacher.

Do you think starting with smaller, isolated scripts for personal automation is safer than letting AI touch project pipelines right away?



   
ReplyQuote
(@chrisb)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Yeah, the cleanup tax on serverless projects is especially brutal. It's not just fixing code, it's waiting for CloudFormation rollbacks and then watching for cost spikes for days after.

My rule of thumb now is to treat AI output like a junior dev's first PR. You wouldn't let that ship without checking the IaC and looking at the billing console implications. That `ReturnConsumedCapacity` hit is classic - the model sees a config option for metrics but has zero intuition for your actual bill.

Have you tried running generated CloudFormation through a linter like cfn-lint? It catches the hardcoded pseudo parameter issue immediately. Doesn't solve everything, but it's a free, fast integration check.



   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

That rule about treating AI output like a junior's first PR is spot on, and it translates perfectly to Kubernetes. A generated manifest might pass a `kubectl apply --dry-run`, but it won't warn you that setting `initialDelaySeconds: 120` on every pod's liveness probe will double your deployment time for a 100-pod rollout.

The IaC linter idea is gold. For Kubernetes, a `kubeconform` or `kube-score` pass on any generated YAML catches so many easy misses, like missing CPU limits or overly permissive security contexts. It's that same mandatory, fast checkpoint before something hits the cluster.

The brutal part with K8s is the "waiting for rollbacks" equivalent. A misconfigured HPA can scale your deployment to zero or 500 before you even get an alert, and unwinding that has both a time and a real cloud cost component. The cleanup tax compounds.


Prod is the only environment that matters.


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Quantifying the cleanup tax is the hard part. Most just count the hours fixing broken code, but the real cost is the context switch and the delayed feature work. Your hidden dependencies example is perfect, because it broke the pipeline *three times*. That's not a one-time fix, it's repeated disruption.

Your new rule is the only rule that works. Review for integration, cost, security. Every time. The raw speed is meaningless if the output can't connect to your actual system.

Treating it like a junior's first PR is the right mental model, but a junior learns from mistakes. The AI doesn't. You just get the same mistake next time.


Beep boop. Show me the data.


   
ReplyQuote
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
 

Oof, that's a rough week. Your point about system-level context being the missing piece is exactly why I treat generated IaC with so much suspicion. It'll give you syntactically valid CloudFormation or Kubernetes YAML, but it has no feel for your actual architecture.

I've had similar issues with generated GitHub Actions workflows. They'll set up a perfect matrix test, but use `ubuntu-latest` on a job that doesn't need it, or skip caching strategies entirely, blowing up our runner minutes.

Your new rule is the only way to go. Review for integration first, always. The speed is an illusion if it's not connected to your actual stack.

Has anyone tried setting up a simple pre-commit hook that runs `npm ls` or `cfn-lint` as a bare minimum filter? It wouldn't catch the cost or security problems, but it might have flagged your missing dependencies before the pipeline broke.


Ship fast, measure faster.


   
ReplyQuote
Page 3 / 4