Skip to content
Notifications
Clear all

AgentGPT after 12 months - honest review from a devops team of 5

26 Posts
26 Users
0 Reactions
19 Views
(@cloud_cost_optimizer)
Honorable Member
Joined: 7 months ago
Posts: 473
 

Your benchmark of ~85% syntax success aligns with our AWS cost optimization work. The critical data point you haven't included is the *rework rate* for that seemingly correct 15%.

We found that syntactically valid Terraform modules often contained logic that violated our reservation strategy, like generating on-demand instances for steady-state workloads. The module would apply cleanly, but the resulting monthly invoice had a 30% premium over what a correctly reserved setup would cost.

The real metric isn't syntax success, but "policy compliance" on first generation. For us, that figure was closer to 40%, turning the promised time savings into a costly audit and refactor process for every generated module.


every dollar counts


   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

Your observation about "incomplete, not wrong" aligns with what we saw in our A/B test of automation generation. The pattern wasn't random, it was systematic.

The agent consistently omitted conditional logic and default field values, especially when they weren't required by the platform's API schema. For example, it would create a full GitLab CI pipeline but leave the `interruptible` flag unset, defaulting to a value that contradicted our team policy. It treated any optional parameter as truly optional, ignoring our internal defaults.

This created a deceptive success metric, because the artifact *worked*, but it wasn't *correct* by our operational standards. The rework to inject those defaults often took longer than building the initial template from a saved snippet.


p-value < 0.05 or bust


   
ReplyQuote
(@gracep)
Reputable Member
Joined: 2 months ago
Posts: 297
 

Not random. It's a gap between API spec and policy. The model sees optional fields and often skips them or picks the first enum value.

We saw it consistently omit tags on AWS resources, defaulting to empty when our policy requires `CostCenter` and `Env`. The JSON schema says tags are optional, so the agent treated them as optional. We had to bake those mandatory defaults into every single prompt.

Your "Daily vs Weekly" is the same class of error. The schema defines both as valid, so it's a coin flip without explicit instruction.


Data over opinions


   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

The "unversioned codebase" point is critical. We hit the same wall.

Our prompts for IAM policy generation broke silently after an AWS Managed Policy update. The model started referencing new, overly permissive actions because they were in the latest schema. No errors, just a sudden compliance violation.

So you're not just maintaining your prompts. You're also tracking and reacting to changes in the target platforms your prompts describe. That's a dependency chain with zero tooling.


Least privilege is not a suggestion.


   
ReplyQuote
(@harukik)
Honorable Member
Joined: 3 months ago
Posts: 400
 

> It needed a human who understood our funnel stages to check it.

That's the part that worries me. If you need a human who already understands the logic to verify the output, what's the real time save? You're just having the AI draft something for an expert to proofread.

How do you even start to build a test case for that kind of fuzzy logic gap?



   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

Exactly. The proofreading step is where the whole efficiency argument collapses. You don't just need an expert to proofread, you need them to run the entire mental simulation the AI should have performed.

Our "test case" was a cost disaster. We tried generating CloudWatch alarms for a multi-tier app. The agent correctly placed alarms on the ELB for 5xx errors and on the RDS instance for CPU. But it completely missed the business logic gap: it didn't tie the ELB's unhealthy host count to the ASG's scaling policy. Syntactically perfect, even conceptually aligned with "monitoring," but operationally useless because the critical escalation path was absent.

You can't write a unit test for missing intuition. The verification requires the same domain knowledge you hoped to offload, plus the additional overhead of auditing for silent failures.



   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

That's such a perfect, painful example. It's like the AI passed the multiple-choice exam but failed the open-book essay on "why?"

Your missing ASG policy link hits the core issue. The agent understood the components, but not the *causal relationships* between them. You can't prompt for intuition you haven't explicitly defined, and defining every single causal chain defeats the whole purpose.

This is why I'm starting to think the real value isn't in *generating* the final artifact, but in using the agent to surface those hidden assumptions in the first place. Make it draft the alarms, then ask it, "What are three ways this monitoring setup could fail silently?" Sometimes its wrong answers reveal the logic we forgot to document.



   
ReplyQuote
(@davidm)
Reputable Member
Joined: 3 months ago
Posts: 270
 

That's a really interesting shift in perspective. Using it to surface hidden assumptions instead of trusting it to generate the final artifact makes a lot of sense.

It reminds me of trying to generate a Docker Compose file for a simple web app. The agent got the services right but completely missed setting resource limits, which is a hard requirement for us. Asking it "what could go wrong with this setup?" probably would have flagged that omission.

So maybe the real workflow is to have it create a first draft as a checklist, then we fill in the intuition it can't have. That feels more honest about the tool's limits. Thanks for this idea.



   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

That's the part I'm struggling with too. You end up reviewing every line anyway, so where's the time save?

But maybe it's not about saving time on the first draft. What if the real value is using the AI to catch your own blind spots? Like, you write a policy, you think it's solid, then you ask the AI to generate three alternative versions. Seeing what it comes up with might show you an edge case you missed completely.

Still, for a simple script or config, it feels faster to just write it from a template. The overhead of prompting and then reviewing everything is real.



   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

You're spot on about it feeling faster to just use a template for simple things. Where I've found it useful is when I'm dealing with a system I don't know inside out, like a new API connector or a cloud service we're adopting.

For example, asking it to draft a dbt model for a new SaaS source. It'll get the basic YAML and SQL structure right, but the real value is in the mistakes - it might join on a field that's not indexed, or use a timestamp that's in UTC when we need local time. Those "wrong" answers force me to check assumptions I didn't even know I was making.

So maybe the time save isn't in the draft itself, but in surfacing those knowledge gaps early. Still, the overhead is real. If I already have a solid template library, I usually just reach for that.


ship it


   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 3 months ago
Posts: 380
 

That 85% success rate for syntactically correct generation in controlled scenarios is the key data point. It aligns with our findings, and it's exactly where the operational friction begins.

The gap between syntax and semantics becomes the primary cost center. For instance, when generating Terraform for a Pub/Sub topic, the syntax will be flawless. But the agent will default the retention duration to 7 days because that's the platform default, not because it's the right value for your pipeline's downstream consumers. You now need a review process to catch all these silent, schema-legal defaults.

This forces you into a new kind of template management: you're not just maintaining prompt libraries, but also a parallel library of context and policy constraints that the API spec doesn't encode. The overhead of maintaining that shadow specification often negates the initial time saved on drafting.


null


   
ReplyQuote
Page 2 / 2