Skip to content
Notifications
Clear all

Hot take: BabyAGI is terrible for long-running tasks.

1 Posts
1 Users
0 Reactions
34 Views
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
Topic starter   [#17741]

Okay, I need to get this off my chest after burning a weekend and a non-trivial amount of credits. I set up BabyAGI to automate a competitive analysis task I do monthly. The idea was solid: let it run, gather data, synthesize, and email me a report. It worked… for about 2 hours.

Then it fell apart. Here’s my breakdown of why it’s fundamentally unsuited for anything longer than a quick, one-off query:

**The Core Problem: Unbounded, Recursive Cost**
BabyAGI’s strength is its loop—create task, execute, enrich list, repeat. That’s also its fatal flaw for long runs. Without a strict, self-enforcing budget or a hard stop mechanism, it will happily:
* Chase edge-case subtasks down rabbit holes.
* Re-research topics it already covered because the context window gets full and it “forgets.”
* Generate new tasks faster than it completes them, leading to exponential growth in API calls.

**My Cost Tracking**
I used a simple sidecar script to track OpenAI token usage. For a 4-hour attempted run:
* First 90 minutes: ~$0.42 (productive)
* Next 150 minutes: ~$3.71 (increasingly repetitive, looping queries)
The cost per minute **tripled** as the task list management degraded. For a true long-running task, this would scale to absurdity.

**The Hidden Pitfall: No Built-in FinOps Controls**
Unlike managed agents platforms, a vanilla BabyAGI script has no:
* Per-session or per-task cost ceilings.
* Auto-halt on diminishing returns (e.g., if last 5 tasks added no new info).
* Awareness of its own running cost. It doesn't “know” to stop or simplify when the bill gets high.

If you absolutely must use it for something extended, you **must** bake in your own circuit breakers. I'm now adding:
* A hard cap on total tasks created (e.g., 100).
* A separate monitor that kills the process if token usage exceeds a threshold.
* A deduplication check for new tasks against the past 20 completed tasks.

Without these, it’s a runaway train. Has anyone else built effective cost-control wrappers for these open-source agent loops? I’d love to compare notes before my next experiment.



   
Quote