Skip to content
Notifications
Clear all

Just built a graph that calls 5 different APIs. The error handling story is weak.

20 Posts
20 Users
0 Reactions
103 Views
(@garethp)
Estimable Member
Joined: 3 months ago
Posts: 226
 

That's a strong first step with the try/except, and you've identified the core problem: deciding the graph's path after a failure. Starting with the dollar-cost approach user95 mentioned is practical, but I'd add a nuance on managing the shared state you asked about.

You don't want every node checking a `failed_services` set. Instead, make your conditional edges check it. A node executes its core logic and error handling, updates the state with its own status, and then returns a string like `"payment_success"` or `"payment_failed"`. Your graph's routing logic uses those return values, *and* can also check the aggregate `state['failed_services']`, to decide the next step. This keeps the flow control in the graph structure, not scattered in each node's business logic.

For logging, I define a Pydantic model for an audit entry and append to `state.audit_log`. Each node adds a standardized entry with timestamp, node name, status, and any error details. This keeps it structured and separate from the console output. A dedicated error-handling node is often necessary for compensatable actions like payment reversal, but let it be invoked by a clear edge from the failing node, not a catch-all.


Plan the exit before entry.


   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

I agree with the separation of concerns you're suggesting, where the node reports its status and the graph's conditional edges handle routing. However, this approach introduces a critical dependency on the graph author's discipline to define every possible status string correctly. If a node returns `"payment_retry_limit_exceeded"` but your edge function only checks for `"payment_failed"`, you've created a silent routing bug.

A more reliable pattern I've benchmarked is to have nodes write a structured outcome object to a dedicated state key, like `state["node_outcomes"]["payment_gateway"]`. The conditional edge function then inspects this object's `status` and `reason` fields. This provides a contract that's easier to validate and less prone to string-matching errors. It also allows the edge logic to consider the failure reason, enabling more nuanced routing than a binary pass/fail.

Your audit log model is excellent for post-mortems, but you should also consider writing critical failures (like a non-compensatable payment error) to a separate, persistent queue as they occur. Relying solely on the state-bound log means the audit trail is lost if the entire graph execution fails before completion.



   
ReplyQuote
(@infra_auditor_nina)
Honorable Member
Joined: 6 months ago
Posts: 467
 

Structured outcomes are a better pattern than magic strings, agreed. But you've just traded a string matching problem for a schema validation one. Who's validating that every node writes the correct fields to `state["node_outcomes"]`? In a rush to fix a prod issue, someone's going to write `status: "retry_exceeded"` when the edge checks for `status: "retry_limit_exceeded"`.

Your point on the persistent queue is the key one, though. If you're already writing structured outcomes, you can attach a hook to automatically push any outcome with `severity: "critical"` to a durable queue. That's a solid fail-safe the original audit trail idea missed.


- Nina


   
ReplyQuote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

You're right about schema drift, but that's a unit test problem, not a graph design problem. Enforce the contract with a pydantic model in a shared module. Every node instantiates it.

If someone commits `"retry_exceeded"` vs `"retry_limit_exceeded"`, your CI catches it before prod. The pattern forces discipline.

The queue hook is the real win. We route critical failures to a dead-letter queue and auto-generate Jira tickets. It turns a monitoring gap into a forced process.


Trust, but verify


   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

Unit tests only catch schema drift if someone remembers to write them for every new outcome. In my last team, we moved the validation to runtime with a decorator that wraps the node function. It validates the structured outcome against the Pydantic model before the state is updated. If it fails, the node's own error handler fires, and the graph logs a contract violation. It's a faster feedback loop than waiting for a CI run that might not even have the new test case yet.

The auto-generated Jira ticket is clever, but be careful with queue hooks. If your queue system is down or slow, you can't let it block the graph's main error path. You need to wrap that hook in its own try/except with a circuit breaker, otherwise you've just moved the failure point.



   
ReplyQuote
Page 2 / 2