Skip to content
Notifications
Clear all

Just built a product feedback summarizer crew - results are okay, but slow.

4 Posts
4 Users
0 Reactions
13 Views
(@elliotn)
Reputable Member
Joined: 3 months ago
Posts: 291
Topic starter   [#26437]

I've spent the last week implementing a CrewAI system designed to process and summarize user feedback from our internal product forums. The primary goal was to automate the weekly digest our product managers currently compile manually. The crew structure is fairly standard: a scraper agent to fetch raw feedback threads, a categorizer agent to tag them by feature area and sentiment, and a summarizer agent to produce the final executive brief.

The quality of the output is acceptable—it correctly identifies top themes and pulls representative quotes with about 85% accuracy compared to my manual benchmark. However, the execution time is problematic. Processing a batch of 50 feedback items takes approximately 12-14 minutes on average. This is far too slow for near-real-time application and even burdensome for a daily batch job.

My crew configuration is as follows:

```python
from crewai import Crew, Agent, Task, Process
from crewai_tools import ScrapeWebsiteTool, FileReadTool
import os

# Define Agents
scraper = Agent(
role='Feedback Collector',
goal='Extract all raw text from specified feedback forum threads.',
backstory='A meticulous data gatherer.',
tools=[ScrapeWebsiteTool()],
verbose=True,
allow_delegation=False
)

categorizer = Agent(
role='Feedback Analyst',
goal='Categorize each feedback item by feature and sentiment, and extract key quotes.',
backstory='An analytical product analyst.',
verbose=True,
allow_delegation=False
)

summarizer = Agent(
role='Summary Writer',
goal='Synthesize categorized feedback into a concise, actionable report for product leadership.',
backstory='A strategic communicator.',
verbose=True,
allow_delegation=False
)

# Define Tasks with explicit async expectations
fetch_task = Task(
description='Fetch all feedback from the forum URLs provided in the context.',
agent=scraper,
expected_output='A list of raw feedback text strings.'
)

categorize_task = Task(
description='Process the raw feedback list. For each item, output: Feature (e.g., "Dashboard", "API"), Sentiment (Positive, Neutral, Negative), and a key quote.',
agent=categorizer,
context=[fetch_task],
expected_output='A structured list of dictionaries containing feature, sentiment, and quote.'
)

summarize_task = Task(
description='Using the categorized list, produce a one-page summary report highlighting top themes, sentiment distribution, and critical user quotes.',
agent=summarizer,
context=[categorize_task],
expected_output='A well-formatted markdown report.'
)

# Form Crew
feedback_crew = Crew(
agents=[scraper, categorizer, summarizer],
tasks=[fetch_task, categorize_task, summarize_task],
process=Process.sequential,
verbose=2
)
```

The bottlenecks I've identified through verbose logging and timing blocks are:
* **Sequential LLM calls:** Each agent's work is a blocking LLM operation. While the process is `sequential`, there's no parallelization *within* tasks (e.g., processing 50 feedback items one-by-one).
* **Tool overhead:** The scraping tool, while convenient, adds latency. The fetched content is often verbose and includes UI boilerplate, which the categorizer LLM call must wade through.
* **Context passing:** The entire raw text output of the scraper (which can be large) is passed into the categorizer task context, potentially increasing token processing time and cost.

I am considering a few optimizations:
1. Pre-processing the scraped data with a simple script to clean HTML/boilerplate before it reaches the categorizer agent.
2. Modifying the crew structure to use hierarchical delegation, where a manager agent splits the batch into smaller chunks for parallel processing by multiple worker agents (though CrewAI's support for this pattern seems nascent).
3. Experimenting with lower-parameter LLM models for the categorization step, which is a relatively simple classification task.

Has anyone else built high-volume summarization or data processing crews? I'm particularly interested in:
* Benchmark times for processing N items.
* Patterns for parallelizing batch processing within a single task.
* Whether using CrewAI's `Process.hierarchical` actually yielded performance gains in practice, or if it's more for organizational logic.
* Any instrumentation or monitoring you've added to track crew latency per task.

The framework is promising for orchestrating complex reasoning, but for data-intensive pipelines, the current execution model feels like it's straining against a batch-processing workload.


Data first, decisions later.


   
Quote
(@charlie2)
Reputable Member
Joined: 3 months ago
Posts: 345
 

That's a neat idea! I've been thinking about something similar for our team's Confluence feedback pages.

> the execution time is problematic

Did you try setting the process to hierarchical? I saw that in a tutorial once. It makes agents work in a stricter sequence, which might cut down on chatter and speed things up. Just a thought!

What would you recommend for getting the accuracy from 85% up a bit higher? Was that mostly about the categorizer agent's instructions?



   
ReplyQuote
(@chrisf)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Have you checked the async options? I'm just starting with CrewAI, but I saw in their docs you can make tasks run concurrently instead of waiting for one agent to fully finish before the next starts. That might shave off some time.

Also, 12-14 minutes feels long. Is it maybe waiting on LLM calls for each tiny item instead of batching them into bigger chunks?

For the accuracy, maybe the categorizer needs clearer examples of what "positive" vs "neutral" looks like in your forums? I've found that helps a ton.


Still learning.


   
ReplyQuote
(@avab)
Reputable Member
Joined: 3 months ago
Posts: 252
 

Hierarchical process can sometimes make things worse, not better. It depends on how much the agents truly depend on each other's output. If they're all hitting the same LLM with sequential calls, you're just adding more orchestration overhead.

And I'd be cautious about chasing a higher accuracy benchmark right now. 85% against a manual check is decent for a first pass. You need to figure out if that 15% error is truly costly or just stylistic differences. Tuning the prompts might get you a few points, but the real cost is the LLM cycles you're burning on every single feedback item.

Before you optimize for time, have you done the math on what this pipeline actually costs per run? That's the number that'll get your finance team's attention faster than the 14-minute runtime.


Question everything


   
ReplyQuote