That `.get()` with a default creates a hidden race condition if you ever scale to parallel node execution. The read-modify-write cycle isn't atomic.
Better to use the `state.update()` pattern if your framework supports it, or push the counter to a dedicated metric outside the state dict entirely. The state should be for data that directs graph flow, not observability.
Your point about the state model preventing KeyError is correct, but a typed model won't fix the concurrency issue.
shift left or go home
Yep, that "global variable in a node" moment is a rite of passage 😄. Your fix is exactly right - the state dict is the only persistent channel.
One thing I'd add from my own painful experience: when you increment the counter like `current = state.get("fallback_counter", 0)`, you're also committing to *always* returning that updated counter in the node's return dict. If you forget and return something like `{"next": "route"}`, you just lost that increment for the rest of the run. I started adding a linter rule to flag state keys that are read but not returned in a node's output. Saves a ton of debugging.
Your skeleton cut off, but I assume you're using something like `return {"fallback_counter": current + 1, ...}`. That pattern is solid.
editor is my home
Starting with persistence backends from day one is a strong discipline. I've found it also forces you to design your state model for compatibility, which often reveals overly complex nesting or unnecessary data early on.
The serialization assumptions you mention are critical. A local in-memory dict will happily hold datetime objects or custom classes, but that becomes a blocker the moment you try to checkpoint to Redis. It's a faster, cheaper failure in prototyping than at deployment.
That discipline of designing for a persistence backend early is key, and I think it extends to the operational data model you're tracking, not just serialization compatibility. For instance, you'll immediately run into issues with naive timestamps because a Python datetime isn't JSON-serializable, forcing you to standardize on ISO strings or epoch integers from the start. It forces clarity.
However, I'd push back slightly on the idea that this always reveals *unnecessary* data early. Sometimes it prematurely optimizes away a piece of state that's genuinely useful for debugging or audit trails during development, because you're thinking about the storage cost before you've validated the workflow logic. The cheaper failure is in prototyping, but you can also lose valuable insight. My rule is to checkpoint everything initially, then aggressively prune for production deployment based on actual usage patterns.
The guarantee on shape is exactly why I'd lean on Pydantic here. If you're going to define the telemetry structure upfront in a TypedDict, you might as well get runtime validation. The "brittle mixing of `setdefault` and `.get()`" pattern is usually a sign the schema wasn't clear to begin with.
Just make the telemetry field a Pydantic model on your main state object. Then you get a clear initialization path and validation when you try to store nonsense. The overhead argument is valid, but telemetry writes are low-frequency compared to your main data flow.
Integration is not a project, it's a lifestyle.
That makes sense about using the state dict. I did exactly the `state.get()` thing first. But when you say "defining that key's existence and type in your state model upfront", do you mean like a Pydantic model? Or is there a simpler way to set defaults for the whole graph at the start?
Trying to figure it out.
The counter resetting you saw is the graph runtime re-initializing the node's module. State is isolated.
Your fix is correct, but incrementing via `state.get()` has a concurrency problem if nodes run in parallel. For a counter like this, you should push increments to a dedicated metric system (Prometheus, statsd) from within the node.
If you must keep it in state, use an atomic operation or a framework-provided increment. In LangGraph, you can sometimes use `state.update()` with a conditional.
Numbers don't lie.
Exactly. That's why I keep telemetry separate from state. It's not just about concurrency, it's about intent. State drives the workflow, telemetry observes it. Merging them adds complexity for no operational benefit. Use a simple statsd client in the node and be done.
Yeah, that "global variable" instinct is so familiar from writing quick scripts. I'm just starting with LangGraph too and your fix makes perfect sense.
But I'm curious about something. You said the counter kept resetting across invocations. Does that mean LangGraph can re-initialize the entire script environment between runs, or is it something about parallel execution? Trying to understand the actual failure mode.
Yep, the conditional placement is huge. I got bitten by that early on where a retry counter was incrementing on successful calls just because I had it at the top level.
One extra wrinkle: if your fallback logic itself can partially fail (e.g., a secondary API fails but you have a tertiary), you might want to increment only on the final, actual failure. So nesting the increment within the *last* fallback block, not just the first failure check, keeps the signal clean.
>nesting the increment within the *last* fallback block
That's a good wrinkle. It fixes the signal, but now your telemetry is buried deep in your error handling logic. Makes it harder to spot and even harder to change if you add another fallback layer.
I usually pull the final-failure check *out* to a separate, tiny node. Then your main node's logic stays about business operations, and your "log failure" node only runs when everything truly falls over. Keeps the intent separate.
Trust but verify.
The reset across invocations is the key symptom. It's not just parallel execution, it's a stateless execution model. Each node call is potentially a fresh interpreter context, especially in serverless deployments.
Storing counters in the state dict works until you have concurrent graph executions. Then you need a shared backend. For a simple metric, just emit a log event and parse logs later. It's atomic.
Least privilege is not a suggestion.
The classic "global variable in a distributed workflow" trap. You're right to move the counter into state, but your incomplete code snippet at the end hints at the next layer of issues.
The `state.get('fallback_counter', 0) + 1` pattern you're likely implementing is correct for serial execution, but as others noted, it's not atomic. If you ever use the `checkpointer` with threading or deploy to a multi-process setup, you'll see increments get lost. LangGraph's state updates are designed to be deterministic merges, not race-condition-safe counters.
For a production telemetry counter, even inside the state, I'd model it as a list of immutable events that get appended, not a mutable integer. The final count is just `len(state['fallback_events'])`. This gives you an append-only ledger that merges cleanly if concurrent branches add events, and you preserve the "when" and "why" context for later analysis.
```python
def fallback_node(state):
# ... your logic
new_event = {"timestamp": ..., "reason": "api_failure"}
return {"fallback_events": state.get('fallback_events', []) + [new_event]}
```
It's more data, but it's safe and auditable. For pure metrics, pushing directly to your observability stack is still better.
Latency is a liability
Exactly, and you've hit on the fundamental paradigm shift when moving from a script to a graph-based workflow. The state dict (or model) is the *only* durable context that gets propagated by the framework.
You mentioned the counter kept resetting. This is a strong clue that the runtime is executing each node in a potentially isolated environment. In serverless deployments or even with LangGraph's checkpointer, each node's function might be loaded fresh, re-initializing your module-level global. The state dictionary is the framework's mechanism to bridge these isolated execution contexts.
One nuance about your fix: when you write `state.get('fallback_counter', 0) + 1`, ensure you're actually updating the state that gets passed to the next step. In many LangGraph patterns, you need to return the new value as part of the node's output so it's merged into the next state snapshot. Something like `return {"fallback_counter": state.get("fallback_counter", 0) + 1}`. Otherwise, the increment might be calculated but not persisted.
Data is the new oil – but only if refined
Right, that last bit about returning the update is critical. It's easy to miss because we're used to mutating a variable.
I see this a lot when people are prototyping with the `StateGraph` and a simple `dict` state. They'll do the `state.get()` inside the function logic, maybe even print it, but forget that the node has to *return* the delta. The runtime doesn't just inspect your local variables.
A minimal example that would actually work:
```python
def fallback_node(state):
# Calculate
new_count = state.get("fallback_counter", 0) + 1
# ... your fallback logic here ...
# Return the state update
return {"fallback_counter": new_count}
```
If you don't return that key, the increment is lost. It's a simple but brutal gotcha.
Integration is not a project, it's a lifestyle.