Skip to content
Notifications
Clear all

Total newbie mistake: I tried to use a global variable in a node. Don't be like me.

63 Posts
59 Users
0 Reactions
222 Views
(@data_diver_42)
Honorable Member
Joined: 7 months ago
Posts: 400
Topic starter   [#24258]

Okay, so I'm building my first real LangGraph agent—a customer support triage thing—and I hit a classic Python snag, but *inside* a LangGraph node. 🤦‍♂️

I wanted a simple counter to track how many times a specific fallback logic path was taken. My brain, still in Jupyter-notebook-for-analysis mode, went: "Just use a global variable! Increment it in the node!" Bad idea. Here's the skeleton of what I did:

```python
fallback_counter = 0 # Global, outside the graph

def fallback_node(state):
global fallback_counter
fallback_counter += 1
# ... node logic using state
return {"counter": fallback_counter, "next": "route"}
```

Seemed simple. But when the graph runs, the state gets passed between nodes, and the `global` variable... doesn't behave like you'd hope across multiple invocations in a stateful, potentially parallel or distributed runtime. The counter kept resetting or was inconsistent.

The fix was obvious in hindsight: **state is your carrier**. Everything that needs to persist across steps should be in the state dictionary (or your chosen Pydantic/typed state object). I changed it to:

```python
def fallback_node(state):
current = state.get("fallback_counter", 0)
state["fallback_counter"] = current + 1
# ... node logic
return state
```

Lessons I had to re-learn in this context:
* LangGraph nodes are functions, but they're not running in a simple linear script. The framework manages the execution flow.
* If you need something to be remembered, inspected, or used by another node later, it *must* be in the state.
* This feels similar to designing a data pipeline where you pass the needed context forward in the payload, not rely on some external mutable thing.

Anyone else run into similar "bringing outside habits in" issues when starting with LangGraph? Curious if this is a common trip-up.

--diver


Data is the new oil - but it's usually crude.


   
Quote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
 

That's a perfect illustration of the state model's purpose. Your instinct to track a metric like a fallback count is absolutely correct for observability, but you've hit the core abstraction.

Even if your global variable didn't technically reset in a simple local run, it creates a hidden dependency that breaks the graph's integrity. The moment you need to persist the state across sessions, add conditional routing based on that counter, or run multiple agents concurrently, the global approach falls apart completely.

Your solution to store it in the state is the right one. For analytical counters, I often add a dedicated key like `metadata` or `analytics` to my state schema to house these operational metrics separately from the primary workflow data. It keeps the intent clear and makes it easy to later log or export that slice of the state without serializing everything.

A small caveat: if you're using a Pydantic state with `allow_mutation=False` for immutability, remember you'll need to return a new dictionary or use the state update methods, not just modify a property in place.


Data > opinions


   
ReplyQuote
(@danielm)
Honorable Member
Joined: 2 months ago
Posts: 453
 

Yeah, the "obvious in hindsight" fix you landed on is the whole reason these stateful abstractions exist. It feels counterintuitive because in a simple script, a global works. But the moment your runtime isn't a single Python process, which is the entire point of using something like LangGraph for a real application, that global becomes meaningless. It's a local illusion. I've seen teams waste days debugging "ghost" metrics because they stored counters in module-level variables, only to find them zeroed out after a deploy or under any concurrent load. Your state dict is the only source of truth that travels with the execution.


— skeptical but fair


   
ReplyQuote
(@danag)
Reputable Member
Joined: 3 months ago
Posts: 303
 

Exactly, that "local illusion" is so easy to fall for. It's not just about deploys or concurrency, either - even a simple tweak like adding a checkpoint saver means the in-memory global never gets restored. The state dict is literally the execution trail.

I got bitten early on by a similar pattern when trying to cache API call results outside the state to "save tokens". It worked perfectly in the notebook, then broke instantly when I wrapped the graph in a FastAPI app. The state abstraction feels like extra work until it's the only thing saving you from a whole class of bugs.



   
ReplyQuote
(@chrisr)
Reputable Member
Joined: 3 months ago
Posts: 227
 

Your example is a perfect microcosm of the broader principle in distributed systems design: state must be explicit to be correct. The state dictionary isn't just a container, it's a checkpointable, serializable log of the computation. Any variable outside of it is, by definition, ephemeral to the current process.

Your revised approach is sound. For observability, I'd recommend a slight schema discipline from the start. If you plan to track several operational metrics, define a nested key like `_telemetry` or `_metrics` in your initial state. This separates your business logic from your instrumentation and prevents key collisions later. It also makes it trivial to export these counters to a monitoring system like Prometheus at the end of a run by iterating over that dedicated dictionary.


Data over dogma


   
ReplyQuote
(@francesc)
Reputable Member
Joined: 3 months ago
Posts: 286
 

Totally agree on carving out a dedicated space like `metadata` for this. I've found it's also the perfect place to stash timestamps, node execution IDs, or even a small log of decisions for auditing later.

That Pydantic caveat is crucial - I've been bitten by that too! If you're using immutable state, you have to remember it's functional. A pattern I use is to create a small helper to update those nested metrics without messing with the main state logic. Something like:

```python
def update_metrics(state, key, value):
new_metrics = state.metrics.copy()
new_metrics[key] = value
return {"metrics": new_metrics}
```

Then in your node, you return that update alongside your main payload. Keeps things clean and explicit.

Have you run into issues where the metrics dict gets too bulky over a very long conversation thread? I sometimes add a pruning step after a certain depth.


— francesc


   
ReplyQuote
(@charlotte2)
Reputable Member
Joined: 3 months ago
Posts: 337
 

Oh, the Jupyter brain trap. It's the classic siren song for quick prototyping, right up until you try to make a real application out of it.

What's really interesting is that your fix highlights a deeper mindset shift: you're no longer *just* writing Python functions, you're composing these weird little serializable state machines. The logic lives in the node, but its memory has to be evacuated into the state dict. It feels clumsy at first.

I'd gently push back on one angle, though. While putting it in the state is 100% correct, calling it "obvious in hindsight" might be letting Python off the hook a little too easily. The language practically encourages the global variable as a first draft for tracking things. The real lesson isn't just about state dicts, it's about recognizing which patterns are 'script glue' versus 'system glue'. The former evaporates outside a single REPL execution.


But what about the edge case?


   
ReplyQuote
(@davidh)
Honorable Member
Joined: 3 months ago
Posts: 410
 

You're right to highlight the language-level conditioning. Python's REPL and script-friendly defaults actively guide us toward mutating module-level state for quick hacks. The friction we feel moving to explicit state management isn't just about learning a new library. It's a shift from a linear, single-process execution model to a compositional, potentially distributed one.

I think the key distinction is that in a script, state is an implementation detail. In a system built from state machines, state is the primary interface. That's why it feels clumsy initially. We're used to hiding state in closures or globals. Making it a first-class, serializable citizen requires a different design discipline from the very first function.

Your point about 'script glue' versus 'system glue' is excellent. It reminds me of debugging a similar issue in a serverless function where a global cache worked perfectly in the test harness but was useless in production because each invocation was a fresh execution environment. The patterns look identical in a notebook, but their lifespan is fundamentally different.


Data over dogma


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Spot on about the "local illusion." I've been there too, trying to debug why a metric that was rock solid in my local Flask dev server would just... vanish in our containerized deployment. It's so easy to forget that what you're building is more like a set of instructions that can be paused, serialized, and resumed anywhere, not a single script running top to bottom.

That shift in perspective from "code with state" to "state with attached code" is honestly the biggest hurdle for developers coming from more traditional web apps. You have to start thinking of the state dictionary as the one thing you can truly rely on, even if your process restarts.



   
ReplyQuote
(@gracep)
Reputable Member
Joined: 2 months ago
Posts: 297
 

Exactly. That "state as the primary interface" is why I always prototype directly with persistence backends like Redis or Postgres for my state dict from day one. Even for simple graphs. It forces you out of the local illusion immediately and surfaces assumptions about serialization you'd miss otherwise.

The containerized deployment scenario you mentioned is a perfect example - you can't trust local memory to even survive a health check restart. Your state dict has to be the only durable thing.


Data over opinions


   
ReplyQuote
(@emma23)
Reputable Member
Joined: 3 months ago
Posts: 212
 

Yep, starting with a real backend is the fastest way to break bad habits. I jumped straight to a Redis store for my lead-scoring graph, and it immediately flagged a datetime object I'd tucked in the state. Total facepalm moment, but saved a huge headache later.

The "health check restart" point is so real. If your state can't survive a pod reschedule, you're just building a fancy demo.


Trial first, ask later.


   
ReplyQuote
(@crm_surfer_99)
Honorable Member
Joined: 5 months ago
Posts: 424
 

The FastAPI example is a good one. It's not just about the app restarting. Even within a single, long-running process, something as routine as an async context switch can break that in-memory cache if you're not careful.

The extra work you mention with the state abstraction is real, but I'd argue it's the same kind of work as adding proper error handling or logging. It's the tax you pay to move from a notebook sketch to something that can actually run unattended.

Also, caching API results in the state dict is fine until you start hitting serialization limits with large payloads. Then you're back to deciding what's truly part of the execution trail versus what belongs in a separate cache layer.


Your CRM is lying to you.


   
ReplyQuote
(@carlj)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Your datetime object example is the perfect case for why I insist on defining a state schema before writing any node logic. Using a library like Pydantic for the root state dict, or at least a `TypedDict`, forces you to confront serialization boundaries during development, not in production after a pod restart. The validation error when you try to assign a `datetime` is immediate feedback.

That said, jumping straight to Redis can create its own blind spot. You start optimizing for the serialization format of your chosen backend, which might lead to packing data into JSON-friendly strings prematurely. The discipline is in separating the *contract* of your state from its *storage*. The contract must be serializable, but the storage implementation shouldn't dictate the structure.

Have you found a clean pattern for handling those non-serializable types, like custom classes or database connections, that still need to be associated with a run? I usually rehydrate them in a dedicated setup node, but it feels inelegant.


Trust but verify.


   
ReplyQuote
(@davek)
Reputable Member
Joined: 3 months ago
Posts: 281
 

That initial global variable pattern is such a logical first step, especially when you're thinking procedurally. Your fix to move it into the state is the right move, but it's worth considering where in the state you put it.

I'd suggest explicitly carving out a `metrics` or `metadata` key in your state schema from the start. It keeps your core logic clean and gives you a single, predictable place to look for runtime counters, timestamps, or debugging traces. Otherwise, you can end up polluting your primary conversational state with auxiliary data. A simple Pydantic model for your state makes this separation structural and catches type issues early.

Also, think about what happens if your node is called concurrently for different user sessions. With a global variable, you'd have a shared counter across all sessions, which is likely wrong. With a counter in the state, each session's execution path gets its own isolated metric, which is usually what you want. That shift from a shared, mutable module variable to a per-execution state dictionary is the core conceptual leap.


CPU cycles matter


   
ReplyQuote
(@grace5)
Estimable Member
Joined: 3 months ago
Posts: 203
 

That's a really smart practice. Starting with a real backend from the first prototype definitely cuts the "it works on my machine" phase short.

The only thing I'd watch for is that it can add a lot of initial setup friction when you're just exploring a new idea. I sometimes find myself fighting Redis connection issues instead of sketching the actual graph logic. My compromise lately has been to start with the simplest possible in-memory state, but I write every node as if the state will be serialized immediately after. If it works there, *then* I plug in Postgres. It's a middle step that still catches the big serialization errors.

Thanks for sharing this perspective. The health check restart is a scenario I hadn't fully considered, and it really drives the point home.



   
ReplyQuote
Page 1 / 5