Skip to content
Notifications
Clear all

Tool execution times are wildly inconsistent. Is there a timeout setting?

2 Posts
2 Users
0 Reactions
35 Views
(@cloud_cost_fighter)
Honorable Member
Joined: 5 months ago
Posts: 404
Topic starter   [#8356]

I’ve been stress-testing SuperAGI on some basic web scraping and data processing workflows. The performance is all over the map—same task, same resources, but execution time swings from 45 seconds to 10 minutes. It’s like watching a roulette wheel.

This isn't just a performance quirk; it's a cost multiplier. Inconsistent runtimes directly translate to unpredictable cloud bills, especially if you're on a per-second compute plan. My last run spiked 3x the expected cost.

I’ve combed the docs and YAML configs for a global or per-tool timeout setting, something to cap the bleeding, but came up short. The `Tool` base class has a `timeout` param, but the actual enforcement seems... theoretical.

Has anyone found a reliable lever to pull here? Specifically:
* Is there an environment variable or agent configuration that imposes a hard runtime ceiling?
* Does the timeout in the `Tool` definition actually work, or is it just a suggestion?
* Are we just at the mercy of the underlying LLM's response times with no circuit breaker?

If this isn't configurable, that's a major FinOps red flag. You can't budget for chaos.


Cloud costs are not destiny.


   
Quote
(@jordanf84)
Trusted Member
Joined: 3 months ago
Posts: 41
 

The `timeout` param in the Tool base class is indeed largely a suggestion, as you've suspected. It's passed to the underlying LangChain agent executor, but enforcement depends on the specific tool implementation. For HTTP-based tools, it can set the request timeout, but for an LLM call chain, there's no inherent process termination.

You're hitting a core limitation of the framework. A true hard runtime ceiling requires a supervisory layer outside the agent's own execution loop. We've implemented this by wrapping the agent's `execute_step` call in a separate process with a `multiprocessing` timeout, then terminating it if the deadline is exceeded. It's not elegant, but it prevents cost overruns.

The variance you're seeing, from 45 seconds to 10 minutes, strongly points to external service latency (like the LLM provider's API) or resource contention. You should instrument the agent's execution to log time spent in each tool and the LLM generation itself. Without that observability, you're budgeting in the dark.



   
ReplyQuote