Skip to content
Notifications
Clear all

Engineering lead: How does it handle retries and fallbacks on tool failures?

2 Posts
2 Users
0 Reactions
22 Views
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
Topic starter   [#10996]

Having recently completed a comprehensive evaluation of SuperAGI's orchestration layer for a critical analytics pipeline project, I found its approach to error handling and resilience to be both sophisticated and, in some aspects, requiring careful configuration. The core mechanisms for retries and fallbacks are managed primarily through the `Agent` configuration and the underlying `Tool` classes, rather than a single global setting.

The system operates on a hierarchical model:

* **Tool-Level Retries:** Individual tools can be wrapped with custom error handling. The framework allows you to define `max_retry_attempts` and a `retry_interval` within the tool's execution logic. Crucially, this is often implemented by catching specific exceptions (e.g., `APIConnectionError`, `RateLimitError`) and implementing a backoff strategy.
* **Agent-Level Workflow Control:** The agent's execution loop handles tool failures based on the returned `ToolResponse`. A `ToolResponse` with `status` set to `"ERROR"` and a descriptive `result` string triggers the agent's reasoning for a retry or fallback. The agent's `max_iterations` parameter indirectly limits the total number of retry cycles possible within a single task.

A practical example from our testing, where we integrated a data extraction tool prone to transient network failures:

```python
from superagi.tools.base_tool import BaseTool
from superagi.lib.logger import logger
import time
import requests

class UnstableAPITool(BaseTool):
name = "Unstable Data Fetcher"
description = "Fetches data from a flaky API endpoint"

def _execute(self, endpoint: str):
max_retries = 3
backoff_sec = 2

for attempt in range(max_retries):
try:
response = requests.get(endpoint, timeout=30)
response.raise_for_status()
return response.json()
except requests.exceptions.RequestException as e:
logger.warning(f"Attempt {attempt + 1} failed: {e}")
if attempt < max_retries - 1:
time.sleep(backoff_sec * (attempt + 1))
else:
return ToolResponse(status="ERROR", result=f"Failed after {max_retries} retries: {e}")

return ToolResponse(status="ERROR", result="Unexpected execution path")
```

**Key Observations and Pitfalls:**

* **The distinction between a "retry" and a "fallback" is agent-reasoned.** The agent, upon receiving an error, may decide to retry the same tool with modified parameters (if your prompt history provides context) or select a different, semantically similar tool from its toolkit as a fallback. This is not automatic; it depends on the agent's LLM-driven planning.
* **State preservation during retries is partial.** The agent's overarching goal and task description remain, but nuanced context from within the failed tool's execution can be lost unless explicitly passed via the `result` field or managed in the agent's memory.
* **Cascading failures require explicit design.** If Tool B depends on Tool A's output, and Tool A fails after retries, Tool B may still be invoked with an empty or malformed input unless you structure the agent's task list to validate intermediate outcomes. We implemented a checkpoint pattern using custom `ToolResponse` statuses (e.g., `"ERROR_CRITICAL"`) to halt the agent's iteration.

For complex workflows, we ultimately built a wrapper `SupervisorAgent` that managed the retry and fallback logic at a higher level, using SuperAGI's `Agent` as an execution unit. This provided more granular control over logs and state persistence across failure cycles than relying solely on the built-in, tool-level mechanisms.

— Amanda


Data > opinions


   
Quote
(@joshuaa)
Trusted Member
Joined: 3 months ago
Posts: 45
 

You're right about the hierarchical model being key. That `ToolResponse` status pattern is actually quite similar to how gRPC interceptors work for service-level retries - you get a clear status code back and the caller decides what to do.

I've found the `max_iterations` limit to be a bit of a hidden gotcha. If an agent gets stuck in a retry loop with a misbehaving tool, it can burn through its entire iteration budget without making forward progress. You really need to pair it with specific tool timeouts.

Have you looked at whether failed tool executions still count against those `max_iterations`? I saw some inconsistent behavior there in early versions.


Design for failure.


   
ReplyQuote