Okay, I need to get this off my chest after burning a weekend and a non-trivial amount of credits. I set up BabyAGI to automate a competitive analysis task I do monthly. The idea was solid: let it run, gather data, synthesize, and email me a report. It worked… for about 2 hours.
Then it fell apart. Here’s my breakdown of why it’s fundamentally unsuited for anything longer than a quick, one-off query:
**The Core Problem: Unbounded, Recursive Cost**
BabyAGI’s strength is its loop—create task, execute, enrich list, repeat. That’s also its fatal flaw for long runs. Without a strict, self-enforcing budget or a hard stop mechanism, it will happily:
* Chase edge-case subtasks down rabbit holes.
* Re-research topics it already covered because the context window gets full and it “forgets.”
* Generate new tasks faster than it completes them, leading to exponential growth in API calls.
**My Cost Tracking**
I used a simple sidecar script to track OpenAI token usage. For a 4-hour attempted run:
* First 90 minutes: ~$0.42 (productive)
* Next 150 minutes: ~$3.71 (increasingly repetitive, looping queries)
The cost per minute **tripled** as the task list management degraded. For a true long-running task, this would scale to absurdity.
**The Hidden Pitfall: No Built-in FinOps Controls**
Unlike managed agents platforms, a vanilla BabyAGI script has no:
* Per-session or per-task cost ceilings.
* Auto-halt on diminishing returns (e.g., if last 5 tasks added no new info).
* Awareness of its own running cost. It doesn't “know” to stop or simplify when the bill gets high.
If you absolutely must use it for something extended, you **must** bake in your own circuit breakers. I'm now adding:
* A hard cap on total tasks created (e.g., 100).
* A separate monitor that kills the process if token usage exceeds a threshold.
* A deduplication check for new tasks against the past 20 completed tasks.
Without these, it’s a runaway train. Has anyone else built effective cost-control wrappers for these open-source agent loops? I’d love to compare notes before my next experiment.