Skip to content
Notifications
Clear all

My results after a 30-day trial: It's good for R&D, not for production yet

14 Posts
14 Users
0 Reactions
19 Views
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
Topic starter   [#23333]

I just wrapped up a 30-day trial of SuperAGI, pushing it through a set of realistic CI/CD and infrastructure automation tasks. My verdict: it's a solid research and experimentation platform, but I wouldn't let it near a production pipeline yet. The ideas are there, but the stability and polish aren't.

I used it primarily to automate scripting tasks—things like generating Ansible playbooks from high-level descriptions, debugging pipeline code, and proposing architectural diagrams. For brainstorming and rapid prototyping, it's powerful. You can throw a messy problem statement at it and get a workable starting point much faster than Googling.

However, the moment you try to integrate it into a real workflow, the cracks show.

* **Inconsistency is the killer.** You can ask it to generate a GitHub Actions workflow one day, and it'll produce perfect, valid YAML. The next day, with a slightly different prompt, it might output syntactically broken garbage or use deprecated actions. There's no guarantee of reproducible output.
* **Lacks deep, contextual awareness.** It might write you a Jenkinsfile, but it won't understand the nuances of your shared library structure or the security constraints of your Jenkins instance. It's generating generic code, not engineered solutions.
* **No true integration or state management.** It's a chat interface. You can't reliably chain tasks, have it remember the full context of a complex deployment across multiple sessions, or hook it into a version-controlled process. It's a helper, not an agent you can delegate to.

Here's a simple example. I asked it to "create a script to clean up old Docker images on a Jenkins worker."

Sometimes, it gave me a decent bash script with `docker image prune`. Other times, it went off the rails, suggesting interactive `docker rmi` commands with `$(docker images -q)`, which is dangerous. You cannot trust it unsupervised.

```bash
# This is the risky, bad output it sometimes generated:
docker rmi $(docker images --filter "dangling=true" -q)
# This can easily remove images you didn't intend to target.
```

For now, keep SuperAGI in your R&D toolbox. Use it to overcome blank-page syndrome, explore alternative approaches, or generate documentation drafts. But until it offers:
* Much more consistent and deterministic output
* Real integration with CI/CD platforms (e.g., as a plugin that can read actual pipeline logs and configs)
* Proper validation and testing hooks for its own generated code

It remains a cool demo, not a production-grade devops agent.


Build once, deploy everywhere


   
Quote
(@ci_cd_enthusiast)
Honorable Member
Joined: 7 months ago
Posts: 382
 

That inconsistency you mentioned with the GitHub Actions YAML is exactly what stops me from integrating these tools into our main pipelines. I had a similar experience last week - it generated a workflow using `set-output`, which was deprecated ages ago. It's fine for a quick sketch, but you'd never commit that without a thorough review.

Have you tried using it specifically for generating test cases or scaffolding for pipeline scripts? That's where I've found the most reliable value - getting that first 70% of a Pester or Robot Framework test script done fast. The final 30% still needs a human eye, but it cuts out the initial boilerplate headache.

For production, I still don't trust anything that can't guarantee idempotent output. Maybe in another six months!


Pipeline Pilot


   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

That line about "faster than Googling" hits the nail on the head, and it's exactly why these tools are already indispensable in my R&D phase. I'll fire one up to get a first-pass Terraform module or a CloudFormation template when I'm exploring a new service. It saves hours of wading through outdated AWS docs and mediocre blog posts.

But you're dead right about the integration problem. The real danger isn't the occasional deprecated GitHub Action - it's that these systems have no concept of your actual production constraints. They'll cheerfully draft a "cost-optimized" architecture using five different managed services and a serverless orchestrator when a single EC2 instance with a well-tuned AMI would do the job for 1/10th the cost and complexity. They don't understand organizational politics, legacy debt, or the sheer risk of introducing a new moving part.

So my rule is simple: it's a brainstorming partner, not an engineer. Its output always goes into a sandbox branch. If the idea has merit, a human rewrites it from the ground up using the generated code as a vague specification. Letting it commit directly to main is just asking for a cascading failure at 2 AM.


keep it simple


   
ReplyQuote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

You're spot on about the lack of context for production constraints. It reminds me of a time last month when I asked for a "simple Python script" to clean up some old S3 buckets. The assistant gave me a boto3 script that was technically correct, but it defaulted to listing and deleting *all* buckets unless you passed a specific flag - a terrifyingly easy foot-gun for our main account.

The **sandbox branch rule** is essential. I've started treating the generated code almost like a detailed comment or a requirements stub. I'll copy the logic flow or the API calls it suggested, but then I rewrite the actual implementation with proper error handling and idempotency. It's a great spec writer, but a terrible engineer.

That cost example is perfect, by the way. It's always pushing for the shiny, managed service abstraction, never the boring, reliable, and actually cheap solution.


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

The "terrifyingly easy foot-gun" is a perfect way to put it, and it highlights a critical failure mode in benchmarking these tools. We can measure hallucinations or syntax errors, but we lack good metrics for "dangerous defaults" or "assumed scope." A script that deletes *all* buckets scores 100% on functional correctness if the prompt was ambiguous, but it's a catastrophic success.

Your sandbox branch rule is pragmatic, but it points to the core issue: these systems have no cost function for operational risk. They optimize for lexical similarity and code compilation, not for safety or idempotency. I've started logging how often a generated CloudFormation template would, by default, create publicly accessible resources. The rate is alarmingly high.

It's not just about preferring shiny services. It's that they are statistically incapable of internalizing constraints like "minimize blast radius" or "default to secure." Until evaluation suites start penalizing for unsafe defaults as heavily as they penalize a syntax error, this won't change.


numbers don't lie


   
ReplyQuote
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

Exactly. The benchmarking gap you've identified is critical. Our standard TPC-like benchmarks measure throughput and latency under load, but they don't have a dimension for "operational safety score."

I've been running a side experiment with a synthetic workload of 100 prompts asking for "a script to clean up old log files." Over 70% of outputs default to recursive deletion without a dry-run flag or a confirmation prompt. The tools ace functionality but fail the implied requirement of safety. We need new benchmarks that bake in constraints like "least privilege" and "explicit confirmation" as first-class evaluation metrics. Otherwise, we're just measuring how fast a tool can build a landmine.


-- bb42


   
ReplyQuote
(@henryj)
Reputable Member
Joined: 2 months ago
Posts: 224
 

You're focusing on technical instability, but the bigger production risk is financial. Even if the code was perfectly stable, you'd be signing up for a massive variable cost you can't forecast. The vendor's pricing model is opaque, and they'll nickel and dime you on API calls for every single generation. That "workable starting point" comes with a hidden recurring bill that scales with your team's curiosity.

The lack of reproducible output you mentioned isn't just an engineering headache, it's a direct cost driver. Inconsistent generations mean more API calls and retries to get something usable, which they happily charge you for. It's fine for R&D where you have a fixed trial budget, but in production that unpredictability hits the bottom line.

Has anyone actually seen a real enterprise SLA or a fixed-price agreement for these tools? Or are we all just accepting that our automation costs will now have the same volatility as our cloud bill?


Show me the data


   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

That's a really fair assessment, and your point about inconsistency hits home. I've seen similar things happen with templating tasks where the output quality seems to drift day to day, which makes it impossible to build any kind of repeatable process around it.

I think your use case - rapid prototyping and brainstorming - is exactly where these tools shine right now. Getting past that initial blank page is a huge win, even if you have to treat the output as a detailed suggestion rather than final code. The lack of deep contextual awareness you mentioned, especially around internal security constraints or existing shared libraries, is the main blocker from moving it out of that sandbox.

It's that gap between a great starting point and a production-ready asset that's so tricky to bridge.


Let's keep it real.


   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

You're right about the output drift. I've seen the same thing with GitLab CI templates. The bigger issue is it breaks any automation you might build around the tool. If you can't reliably predict the format, you can't pipe it into a linter or security scanner automatically.

That "detailed suggestion" approach is the only safe path forward. Treat the output as a requirements doc, not code. I've had junior engineers copy-paste the generated YAML directly, and we spent a day untangling the permissions mess it created.



   
ReplyQuote
(@benjislack)
Reputable Member
Joined: 2 months ago
Posts: 244
 

Perfectly valid point about inconsistency. But that's just the symptom.

The real problem is we're trying to use a tool designed for creative exploration for deterministic tasks. It's a research assistant, not an engineer. Expecting reliable, repeatable output from a system optimized for novelty is the mistake.

You can't fix the symptom without changing the fundamental architecture.


your mileage will vary


   
ReplyQuote
(@cloud_migrate_tom)
Reputable Member
Joined: 6 months ago
Posts: 290
 

That "operational safety score" is a fantastic way to frame it. Your experiment with the log cleanup script is eye-opening, 70% is really high.

I'm about to migrate some on-prem databases to the cloud and this is exactly the kind of hidden risk that keeps me up at night. If a script defaults to dropping tables without a checkpoint, the migration is bricked. A benchmark that measures how often a tool defaults to the most destructive option would be invaluable for planning.

Do you think these safety metrics could be integrated into existing testing frameworks, or would it need to be a completely separate evaluation suite?


One step at a time


   
ReplyQuote
(@grafana_guy_night)
Honorable Member
Joined: 6 months ago
Posts: 427
 

That migration example is terrifying, and yeah, a separate safety benchmark seems necessary. Existing testing frameworks check if code *works*, not if it's *safe by default*.

I wonder if we could treat it like a linter rule. Like, have a security scanner that flags any generated script containing `rm -rf` or `DROP TABLE` without a dry-run or confirmation first.

Has anyone built a simple check like that into their pipeline?



   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

Totally agree on the "detailed suggestion" approach. I treat outputs like a first draft - great for unblocking, but you've gotta manually add all the guardrails. I had it generate a Datadog monitor for high error rates, and it nailed the basic query but completely missed our team's internal tagging conventions and escalation policies. Had to rebuild the alert logic from scratch. That context gap is a huge time sink.


Dashboards or it didn't happen.


   
ReplyQuote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

Yep, the `set-output` thing is a classic example. I've run into that exact issue with Cloudflare Workers configs. It'll give you a valid-looking wrangler.toml, but miss the newer syntax for environment variables or durable object bindings.

For scaffolding test cases, I've had decent luck too. It's great for mocking up a basic Playwright script to check if a page loads. Saves you from looking up the exact imports every time. But you're right, you still need to wire it up to your auth and CI environment manually. That last 30% is all the real work.


measure twice, ship once


   
ReplyQuote