Skip to content
Notifications
Clear all

Help: BabyAGI outputs are becoming repetitive and low quality.

2 Posts
2 Users
0 Reactions
14 Views
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
Topic starter   [#5740]

Hey everyone, I've been putting BabyAGI through its paces on some automation tasks—mostly around generating CI/CD pipeline scripts and organizing project backlogs. Lately, I'm hitting a wall where the outputs feel like they're on a loop.

The agent starts strong, but after a few iterations, its tasks and suggestions become repetitive. For example, when I ask it to break down a feature into sub-tasks, it keeps rephrasing the same three basic steps instead of diving deeper or considering edge cases. The quality degrades from "useful draft" to "vague placeholder."

I'm running a fairly standard setup with OpenAI and a Pinecone vector store. My hypothesis is the issue might be in the task creation or prioritization steps. Has anyone else experienced this and found tweaks that help?

A few things I've already tried:
- Experimenting with different `OBJECTIVE` phrasing to be more or less specific.
- Adjusting the number of tasks returned per cycle.
- Swapping the LLM from GPT-4 to GPT-3.5-turbo (which made it worse, unsurprisingly).

I'm curious if the community has run into similar patterns. Specifically:
- Are there prompt modifications for the task-creation agent that encourage more variety?
- Does the context window of the vector store significantly impact this repetition if it gets clogged with similar entries?
- Would a different task-execution or prioritization loop (like a custom agent) break the cycle?

I love the framework's potential, but hitting this repetition bottleneck is frustrating. Any insights or shared experiences would be awesome.


Ship fast, measure faster.


   
Quote
(@crm_hopper_2025_new)
Honorable Member
Joined: 4 months ago
Posts: 365
 

Yeah, that's the classic BabyAGI death spiral. Tweaking the objective phrasing never fixed it for me either.

The core problem is usually the task list itself becoming an echo chamber. Each new task is generated from the context of the previous, increasingly generic outputs. You end up with a shallow pool of stored results that just get re-queried.

What finally helped me was hacking the prioritization prompt to aggressively deprioritize any task that sounded semantically similar to the last three completed tasks. It's a band-aid, but it forces the agent to look for a different angle, at least for a few more cycles. After that, it usually collapses again.

Honestly, I moved on. These toy agents are great for a demo, but for actual project breakdown I get better, less loopy results from a structured prompt in a plain ol' ChatGPT session.



   
ReplyQuote