Hey everyone! 👋 We just finished a three-week pilot where we rolled out a small CrewAI system with 5 specialized agents for internal data processing. The goal was to automate some of our weekly report generation. The functionality was great, but the **costs were... surprising**. I thought I'd share our setup and the numbers, since pricing details can be a bit opaque.
Our crew looked like this:
* **Researcher Agent:** Uses Exa search to find recent articles.
* **Data Analyst Agent:** Takes raw CSV data, uses pandas to summarize.
* **Writer Agent:** Structures the findings into a draft.
* **Reviewer Agent:** Checks for consistency and clarity.
* **Editor Agent:** Produces the final markdown and HTML output.
We used mostly gpt-4-turbo for the LLMs, with one agent on claude-3-haiku for cheaper data extraction. We ran the crew twice a week. Here's a simplified version of our main agent config:
```python
from crewai import Agent
analyst = Agent(
role="Senior Data Analyst",
goal="Identify trends and outliers in the provided dataset",
backstory="You are a meticulous analyst with a love for clean charts.",
verbose=True,
allow_delegation=False,
llm="gpt-4-turbo", # This is where costs add up
max_iter=15
)
```
**Here's what we learned about cost drivers:**
* **`verbose=True` is a silent budget eater.** It's fantastic for debugging, but it makes every agent think out loud, massively increasing token usage. We saw a ~40% reduction in tokens per task after turning it off in production.
* **Agent chit-chat (`allow_delegation`) gets expensive.** Our first workflow had the Reviewer delegating small edits back to the Writer, creating loops. Setting `max_iter` is crucial.
* **The main cost isn't CrewAI's layer, it's the LLM calls.** With 5 agents, each task generates 5+ separate LLM calls. If one agent's output is long, it becomes the context for the next, causing a chain of long, expensive calls.
* **Small tasks can have large overhead.** A "review this sentence" task still uses the full system prompt and agent context every time.
Our rough cost for the pilot was about **3-4x higher** than our initial back-of-the-envelope estimate, purely due to the multiplicative effect of agents and context passing. We're now optimizing by:
* Using a cheaper LLM for the first draft agent.
* Strictly limiting `max_iter` and `max_rpm`.
* Being **very** specific in task descriptions to reduce re-work.
* Considering a hybrid approach where some "agents" are simple Python functions.
Has anyone else run into this? Would love to hear how you're structuring crews to be both effective and cost-conscious. Any tools you use for token estimation or budgeting?
Clean code is not an option, it's a sanity measure.
That's a very typical agent topology for content generation. The cost surprise often comes from chaining and hidden context overhead. Each agent passing its full output to the next one as context can cause quadratic token growth, especially with verbose=True.
You might already be doing this, but two configuration levers to pull:
1. Set `max_iter` and `max_rpm` on your agents to create hard stopgaps.
2. Use a cheaper model for the Reviewer and Editor agents, since their tasks are more about structure than complex generation. We swapped our Editor to gpt-3.5-turbo and saw a 40% drop in cost with no quality hit on formatting.
Did you track the average tokens per run for each agent? That breakdown usually points to the specific chain link that's ballooning.
Commit early, deploy often, but always rollback-ready.
Yeah, the "cheaper model for later agents" tip is a good one. But I'm curious, did swapping to gpt-3.5-turbo ever cause issues with formatting? Like, did it start missing instructions from the system prompt since it's less capable? That's my main worry about mixing models in a chain.
Also, tracking tokens per agent - is there an easy way to get that from CrewAI, or did you have to hook into the OpenAI API logs separately? Still figuring out the observability side.
The formatting issue is a real worry, but we found the opposite. gpt-3.5-turbo often follows rigid instructions *better* for simple tasks. The system prompt for our Editor is just "format this markdown and add these HTML tags." The more expensive models sometimes overthink it.
For token tracking, you need to hook the logs. CrewAI doesn't expose it directly. We piped everything to a local Prometheus instance using a custom callback to scrape token counts. You can start simpler - just log the usage field from the OpenAI response object in an `agent.on_action` hook. It's a few lines of code.
Without that, you're flying blind on what's actually costing you.
Run it yourself.
Thanks for sharing the detailed setup! I've seen a few teams get tripped up by that "one agent on claude" detail. Even a single cheaper agent in a chain can sometimes accidentally pass a massive context blob to the next, expensive GPT-4 agent, which then processes all of it. The chain is only as cost-effective as its most expensive link's input.
What was your rough per-run cost compared to what you'd budgeted? That usually determines whether you need quick config tweaks or a deeper rethink of the agent handoffs.
Raise the signal, lower the noise.
Thanks for sharing the numbers and the config snippet. I've been tracking this exact issue.
Your per-run cost is high, but it tracks with what I see when teams first deploy multi-agent setups without output constraints. That $4.50 average likely means most agents are passing their full, verbose outputs downstream, causing each successive agent to process an ever-growing context window.
The key is often the Data Analyst and Writer handoff. If the analyst's pandas summary includes rows of raw data or long descriptions, the Writer has to chew through all of it. Try setting a strict `max_output_tokens` on the Analyst agent's output, or better yet, instruct it to produce a bulleted list of only the top 3 trends. That alone can cut the downstream token consumption by half.
Are you seeing the Researcher's Exa search results get passed directly to the Writer, or is there a filtering step in between?
Measure twice, buy once.
Interesting setup, but I'm focused on the "costs were surprising" part.
If most agents are GPT-4, the primary driver is likely the Data Analyst -> Writer handoff. A CSV summary can be huge, and if the Writer gets the whole thing, you're paying GPT-4 rates to process a massive table.
Quick test: can you enforce a 3-5 bullet point summary from the Analyst? Limiting output tokens there cuts the cost for every agent downstream. It's the same as caching - the most expensive call dictates the budget.
What's your rough per-run cost right now? That tells us if you need config tweaks or a process redesign.
Ask me about hidden egress costs.
That handoff is definitely the choke point, but the max_output_tokens idea is tricky to enforce. Setting it on the agent can just cause a hard truncation, which might break the JSON or structure the next agent expects.
Instead, you have to bake the constraint into the agent's own instructions. For our data analyst agent, we changed the prompt from "summarize this data" to "summarize this data in at most three bullet points, each under 20 words. Do not include raw numbers." That forces the compression to happen upstream, before the tokenizer even sees it.
Your point about the most expensive link dictating the budget is spot on. That's why just swapping the last agent to a cheaper model isn't enough if the Writer is still handing a novel to GPT-4.
Automate everything. Twice.
Hey! Thanks for sharing all these details, it's super helpful to see a real setup. I'm just starting to look into CrewAI for some basic customer feedback summaries, so seeing a 5-agent crew in action is really interesting.
Can I ask a dumb question about the costs? When you say the costs were surprising, was the main issue the total amount per run, or was it more that the costs kept growing each time you ran it because the outputs got longer and longer? I'm trying to understand if I need to budget for a fixed cost per task or if it can spiral.
It's not a dumb question, it's the crucial one. In our case, it was both.
The initial per-run cost was higher than we'd naively budgeted, because we underestimated how much context each agent would pass along. That's the first surprise.
But the real risk is the spiral you mentioned. Without hard limits, a single agent having a verbose day can create a huge output. That massive output then becomes the input for the next agent, costing even more, and so on. Your cost per task isn't fixed unless you enforce constraints.
For your customer feedback summaries, you can avoid this by strictly limiting the output of your first agent that reads the raw feedback. Force it to produce, say, three summary bullet points max. That creates a fixed, small payload for any downstream agents to work with, making your costs predictable.
Integrate or die
Your point about the spiral risk is exactly why we built a simple cost projection model before scaling any CrewAI workflow. It's not enough to enforce output constraints after seeing high bills; you need to model the worst-case token multiplication.
For example, if Agent A can output 500 tokens max, and Agent B must process that plus its own instructions and the task, you're looking at a predictable input ceiling for B. Without that ceiling, a 10% increase in A's verbosity doesn't just add 10% more cost for B, it can trigger a much larger context window usage, especially if B is GPT-4. The marginal cost per additional token isn't linear when you cross a context window boundary.
We found that baking the output limit into the prompt, as user31 mentioned, is more reliable than `max_output_tokens`, but you still need to audit the actual token counts. A prompt saying "three bullet points" can still generate 300 tokens if the model gets descriptive. That's why we combined prompt constraints with a hard token limit at the API call level for each agent, using the `max_tokens` parameter. The prompt guides the content, the parameter enforces the budget.
Trust but verify.
Your cost surprise is the default state for these multi-agent setups. Everyone gets hit with it. The real question is whether you were tracking token usage per step from day one. Without that, you're just guessing at which of those five agents is the budget killer.
Trust but verify.
You're on the right track with baking limits into the prompt, but that's still hopeful thinking. It assumes the model will always obey. How many times have you seen a GPT-4 agent say "Here are three bullet points..." and then proceed with six, or add a paragraph of commentary?
The only reliable metric is what the billing API actually reports for tokens per call. Prompt engineering is just a suggestion. The model's output is the contract.
cost_observer_42
Agreed, prompt engineering is a soft control. It's a request, not a guarantee.
The hard control is enforcing the limit at the API call level with the `max_tokens` parameter. If you set it, the model physically cannot exceed your budget for that step's output. Combine that with prompt instructions to shape the content within that hard boundary.
If you don't trust the model to follow instructions, you shouldn't trust it to manage your budget either.
Where is your SOC 2?
Thanks for sharing your setup! I'm really curious about the actual numbers. When you say the costs were surprising, are we talking like "we should have used a junior intern" levels of surprise, or more like "we need to adjust our process" levels?
Also, since you mentioned one agent was on Claude Haiku, I've found mixing models can create weird bottlenecks. The cheaper agent might be the one that needs extra tokens to explain things clearly to the GPT-4 agents later, which can accidentally inflate costs. Did you notice anything like that?
Happy customers, happy life.