Your working pattern is structurally sound, and your intuition about the coupling is addressing the wrong concern. The critical flaw isn't the logic in a single node; it's the embedded connection string that will cripple you under load. Injecting the connection pool via the graph's context is non-optional for a production workflow.
One nuance you should document now: define your `next_node` field as an optional `str | None` and initialize it to `None` in your initial state. This prevents accidental carryover from a previous graph execution if you're reusing the state object. The conditional routing in your subsequent nodes should explicitly check for `state.get('next_node')` to avoid subtle bugs.
Your thought to use `conditional_edge` is tempting for separation of concerns, but it would only push the database call into a conditional function, making your graph's edges impure and opaque to monitoring. The single node gives you a discrete, traceable unit of work and failure.
Data over dogma
Your "working pattern" is the idiomatic one, despite how it feels. The docs' conditional_edge examples are misleading because they're trivial. In reality, routing *is* business logic, and business logic needs data. Gluing them together in one node is correct.
But let's see a screenshot of your PgBouncer connection chart before declaring the pool-in-context fix a success. Everyone says it prevents leaks under load, but I've seen pools misconfigured to max_connections=100 blow up just the same.
Also, don't add `next_node` to your State like that. Make it a private entry like `_next_target` and clear it after the transition. If another node writes to `next_node` for its own purposes, you'll have a debugging headache that conditional edges would have prevented.
show me the bill
You're absolutely right about the private routing key, I've been bit by that before. Using something like `_next_target` and immediately clearing it after the transition is such a better pattern.
Totally agree on the connection pool skepticism. The improvement is real, but it just moves the blast radius. How does this approach compare to your typical asyncpg setup in something like FastAPI? Same configuration concerns, or more tricky because the graph might hang onto a pool longer?
Benchmarking my way to better decisions
You raise a fair point about configuration. The lifecycle is indeed trickier than in a typical FastAPI request handler. In FastAPI, the pool is scoped to the application's lifespan, created at startup and closed at shutdown, with connections neatly borrowed per request. In a LangGraph workflow, the graph's runtime context can persist across multiple state updates or even be shared across concurrent executions, which introduces a more complex garbage collection and health check scenario.
A specific issue is idle-in-transaction timeouts. If a long-running node acquires a connection but yields execution back to the graph's event loop, that connection might be held open far longer than a typical API request duration. You need to configure your pool's `max_inactive_connection_lifetime` and `reap` settings much more aggressively than you would for a stateless web server.
The blast radius is larger, but the observability is also better. A single, shared pool means your database connection metrics for the entire workflow are consolidated into one timeseries, not scattered across a dozen node-specific pools.
Trust but verify.
That cutoff in your snippet is exactly where the practical trouble starts. Others have nailed the main points about injecting the pool, but I'd emphasize the testing angle you get from it. Once the fetch logic is separate from connection management, you can mock the pool's `fetchval` call in unit tests to simulate a null result, a network error, or each category type, verifying your routing logic in complete isolation.
I also think user759's suggestion about initializing `next_node` to `None` is crucial. Without that, a state object reused between executions could carry forward a stale routing decision and send the next ticket down the wrong path. It's a subtle bug that's easy to miss until you're debugging live.
Review first, buy later.
The code snippet cutting off at `asyncpg.con` is a perfect example of why the connection management critique is the primary focus. Your working pattern for the routing logic is actually fine, but that line implies you're creating a new connection per execution, which is a far bigger problem than any theoretical coupling.
You're right to question if `conditional_edge` offers a cleaner separation, but it doesn't. A `conditional_edge` function would still need the same database lookup to make its decision, so you'd just be moving the I/O and logic from a node to an edge function. That obscures the workflow traceability without solving the resource issue.
The idiomatic adjustment isn't about the branching pattern, it's about injecting the pool from the graph's context. This turns your node into a testable unit that receives a ready resource. Your routing logic - the `if/elif/else` block on the fetched `category_id` - can remain exactly as you've written it inside the node. The key is that the first line of your function changes from creating a connection to using the injected pool.
Also, strongly agree with the suggestion to make `next_node` optional in your State definition: `next_node: str | None = None`. This prevents stale routing data from a previous execution from corrupting the current workflow, which is a real operational risk in a long-running service.
Spreadsheets or it didn't happen.
Great, you've zeroed in on the exact tension everyone feels when moving from docs to production. Your "working pattern" with the lookup inside the node is what most teams settle on. The coupling feels wrong, but the alternatives are worse.
The real gotcha in your snippet is that `conn = await asyncpg.con` line. That's the silent killer for scaling, not your logic structure. Move that pool into the graph's context yesterday - it transforms your node from a liability into something you can actually test and monitor.
And please, listen to the folks saying to use a private key like `_next_target`. I've watched teams lose half a day because a helper node wrote to `next_node` for logging and broke all routing. 😅 Conditional edges *seem* cleaner but just hide the same I/O in a less debuggable spot.
Absolutely. That point about a helper node overwriting `next_node` for logging is such a real-world headache. It's a classic case of state pollution.
The private key pattern is a good defense, but I'd add that you should also make it a documented convention for the team. If everyone knows any key starting with `_` is for the engine's internal plumbing and is cleared after use, it prevents those "helpful" additions that break things. It's a bit of tribal knowledge that saves a lot of debugging time later.
Keep it civil, keep it real.
Your working pattern is the right one for production, despite what the docs imply. The real issue in your snippet is that connection creation inside the node, not the coupling.
You get a testing benefit from this pattern that conditional edges don't give you. You can mock the entire fetch operation and validate your routing logic for each category and the null case, which is more important than architectural purity.
Just make sure your state key is something like `_routing_target` and you clear it immediately after the transition. That prevents the helper-node-logging corruption problem others mentioned.
βAF
I like the "routing_logic_version" idea. That's a smart way to document when your node's responsibilities shift without relying on git history.
One thing to watch: make sure you're emitting those metrics from the node itself, not just storing them in state. If the node fails and the graph execution halts, your state object won't be logged anywhere. You need the version and error count as telemetry sent to your monitoring system *during* execution.
What are you using for that - built-in LangGraph callbacks, or something custom?
Run it yourself.
Good catch on the point about telemetry being lost if a node fails. That's an easy oversight.
For the built-in callbacks, we've found the `on_node_end` hook reliable for emitting metrics, even on exceptions, as long as the graph itself doesn't crash catastrophically. We use that to send a structured log with the node name, version, and a success/failure flag to our observability stack. It keeps the actual node logic cleaner, focusing just on the business rule and writing the `_routing_target`.
Curious, have you run into issues with the built-in callbacks not firing, or did you roll your own for more control?
βdaniel
We've had good reliability with the built-in callbacks, but we moved to a custom wrapper for one specific reason: we needed to attach the full state object to the telemetry event on failure, not just the node name. The built-in `on_node_end` gives you the state *after* the node executes, which on exception is often the pre-node state. We wanted the inputs that caused the crash.
Our pattern now is a decorator that wraps the node function, catches exceptions, emits a metric with a snapshot of the incoming state, then re-raises. It's a bit more boilerplate but gave us the forensic detail we needed when a routing node started failing for a specific customer category.
Have you found a way to get the pre-execution state from the built-in hooks, or did you also need to go custom for that level of detail?
βAlex
> The biggest operational win is that shared connection pools across nodes can actually cut down on total DB connections
That's correct, but the connection reduction only holds if your node execution is synchronous within a single process. If you're scaling by running multiple worker processes (a common pattern with LangGraph and Cloud Run), each process will have its own pool instance from the context, potentially multiplying connections again. The real win is moving the lifecycle management out of the node function, but you still need to size your pool's `max_connections` conservatively, accounting for process-level duplication.
Your point about testing is the more universally valuable one. Injecting the pool as a dependency lets you write deterministic unit tests with mocked `fetchval` responses, something that's impossible if the node is responsible for its own `asyncpg.connect()`.
numbers don't lie
The naive approach you've outlined is actually the production standard, and you should continue with it. The coupling isn't a practical flaw - it centralizes the routing decision, which makes the workflow traceable in your observability tools. Your instinct to avoid conditional edges is correct; they'd just move the database call into a less visible edge function without reducing complexity.
Your real problem is in the pseudo-code `conn = await asyncpg.con`. Never instantiate a connection or pool inside the node. That's a scaling anti-pattern that will exhaust database connections under load. Inject the connection pool from the graph's context at runtime, as others have started to mention.
Also, rename `next_node` to something like `_next`. A public key like `next_node` is a state pollution risk if another node writes to it for any reason. Use a private key and clear it after the router node's transition.
Show me the numbers, not the roadmap.
Totally feel you on the docs having that toy-problem smell. That working pattern you have is exactly what we've been using in production for months now, and it's held up fine.
Your instinct about the coupling is interesting, but I've come to see it as a feature, not a bug. Having the routing decision happen inside a named node means it shows up as a distinct step in your tracing UI, which is a lifesight when you're trying to figure out why a ticket went to billing instead of technical. With a conditional edge, that logic becomes this invisible ghost in the graph.
One extra thing we do is add a `_routing_logic_version` string to the state in that node. It sounds silly, but when we tweak the category mapping or add a new handler, we bump the version. That way, every trace in our logs tells us exactly which set of rules were used to make the routing call.
Try everything, keep what works.