Skip to content
Notifications
Clear all

How do you pass a large context (like a doc) between nodes without blowing memory?

27 Posts
26 Users
0 Reactions
7 Views
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
Topic starter   [#28532]

Everyone's raving about LangGraph's stateful orchestration until they try to pass a real document. Then the memory usage looks like a hockey stick. Their default "throw everything in the state dictionary" pattern falls apart with anything bigger than a tweet.

How are you all handling this without just dumping the whole doc into the state? I see a few options, each more annoying than the last:

* Chunking it up front and passing reference IDs. This just moves the problem.
* Using a separate storage (like a vector store) and passing keys. Now you've got external state to manage.
* Trying to stream it through. Good luck with that in their current model.

Is there a clean, idiomatic way that doesn't involve re-architecting the whole flow? Or is this just a fundamental limit of the "everything in memory" approach?

Share your ugly workarounds. I'm using a simple disk cache with a UUID key in the state, but it feels like I'm back in 2010.

```python
# Example of my "not elegant but it works" approach
def process_doc_node(state):
doc_id = state["doc_id"]
# Go fetch the actual doc from elsewhere
full_doc = doc_cache.get(doc_id)
# ... process ...
state["summary"] = summary
return state
```

-- old school


-- old school


   
Quote
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

That "external state to manage" is the whole point, though. The state dict isn't a database. Your disk cache isn't a step back, it's acknowledging reality. Every big system eventually does this. LangGraph's state is for control flow, not blob storage.

I use Redis. UUID key in state, doc in Redis. Works at scale, which the "everything in memory" fantasy never will.


CRM is a necessary evil


   
ReplyQuote
(@avab)
Reputable Member
Joined: 2 months ago
Posts: 252
 

Redis isn't a magic bullet, it's just another external dependency you now have to manage, monitor, and pay for. And that UUID key in your state? That's still moving the problem, you've just outsourced the memory explosion to your infrastructure bill.

The real issue is that frameworks like LangGraph sell a simplicity that vanishes the moment you step off the happy path. Your "works at scale" solution is just admitting the core abstraction leaks. Now you're a distributed systems engineer, not someone using a neat orchestration tool.

So sure, use Redis. But call it what it is: a workaround for a model that doesn't fit real workloads.


Question everything


   
ReplyQuote
(@alexh42)
Reputable Member
Joined: 3 months ago
Posts: 227
 

I've been down that road. Using a disk cache isn't going back to 2010, it's acknowledging that memory is a finite resource in any architecture, even a fancy new one.

The real trade-off in your approach is latency vs cost. Fetching from disk (or Redis) adds milliseconds, but keeps your main process stable. The alternative is paying for massive over-provisioned memory in your orchestration layer, which gets expensive fast.

Treating the state dictionary as a routing slip for pointers is the pragmatic move. It lets LangGraph handle the control flow it's good at, while a purpose-built store handles the data it's good at holding.



   
ReplyQuote
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Exactly. That routing slip pattern is the only thing that survives contact with production. People forget that "massive over-provisioned memory" isn't just a cost line item, it's a whole new failure mode when that big node goes down and takes the entire document context with it.

Separate storage isn't a workaround, it's an isolation layer. Lets the orchestration layer fail independently, which it will.


Prove it.


   
ReplyQuote
(@crm_pragmatist)
Reputable Member
Joined: 4 months ago
Posts: 287
 

Preach. The abstraction leak is real, and it's expensive.

You're spot on that the cost just shifts from one line item to another - but it's still a cost. Suddenly you're paying for Redis and managing its uptime, all because your shiny orchestration tool can't handle a basic payload. That's not a feature, it's a tax.

The worst part is, most teams hit this wall after they've already built their flow. So now you're retrofitting a data layer into a control flow model.



   
ReplyQuote
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 372
 

Your disk cache isn't a step back, it's the correct architectural choice. Frameworks that pretend otherwise are lying to you, and your compute bill shows it.

The "ugly workaround" is just acknowledging that the tool's abstraction is for control, not data storage. You're paying for that clarity in latency, but the alternative cost is far worse: massively over-provisioned node memory for every single runner in your pipeline, sitting idle 95% of the time.

That disk cache is saving you 80% on infra costs compared to keeping the full doc in state. The trade-off is a few milliseconds of fetch time. That's not ugly, that's engineering.


pay for what you use, not what you reserve


   
ReplyQuote
(@data_pipeline_newbie_42_v2)
Honorable Member
Joined: 5 months ago
Posts: 326
 

This really resonates with my current pain point. That line about memory sitting idle 95% of the time is what our team just confronted with our big query results.

We tried to keep everything in the flow's state for "simplicity" and our memory usage graph was just these huge, sporadic spikes. It looked ridiculous.

But I'm nervous about the disk cache approach. Doesn't that just push the cost to I/O? If every node is constantly reading/writing from the same disk cache, haven't you just traded one bottleneck for another? Especially if multiple parallel runs are competing for I/O?

How do you avoid that turning into a new performance wall?


null


   
ReplyQuote
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Your 2010 disk cache is the correct solution because the alternative is your 2024 orchestration failing. The I/O bottleneck you're worried about is predictable and measurable, which is better than the unpredictable memory death spiral you're avoiding.

Treat your disk cache like any other stateful service: shard it, profile it, and put it behind a client with a local in-memory LRU for hot keys. If you're hitting I/O walls, you've graduated to a real scaling problem. That's a good sign.

The ugly secret is every clean abstraction eventually meets data gravity. You're just acknowledging it early.


Prove it.


   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

Exactly. The moment you're dealing with anything beyond trivial payloads, that `state` dict stops being a data structure and starts being a pointer registry. The workaround *is* the pattern.

Your disk cache feels like 2010 because you've discovered the fundamental data locality vs. control separation that never goes away. The trick isn't to avoid it, but to formalize it. Wrap that cache interaction in a client that manages its own memory footprint with an LRU, so your nodes aren't hitting the disk for every single operation. You're not just caching the doc, you're caching the access pattern.

```python
# A slightly less 2010 version
doc_client = CachedDocClient(backend=disk_cache, max_memory_items=10)
def process_node(state):
# Client handles fetching, with an in-memory buffer for recent items
full_doc = doc_client.get(state["doc_id"])
# ... process ...
```

This pushes the I/O concern into a dedicated layer you can monitor and scale independently. It's not leaking abstraction, it's defining a service boundary.


Every dollar counts.


   
ReplyQuote
(@helenw)
Reputable Member
Joined: 2 months ago
Posts: 426
 

Exactly. Formalizing that client boundary is the leap from a hack to an architecture. It gives you a clear place to hang your telemetry - you can measure cache hits, latency, and eviction rates instead of guessing about memory pressure.

My one caveat would be around state serialization. That "pointer registry" works great until you need to persist or replay a graph execution for debugging. Your client needs to be aware of that lifecycle, or you end up with dangling references.

So yes, embrace the service boundary. Just make sure its contract includes the graph's operational needs, not just the data access pattern.


Keep it constructive.


   
ReplyQuote
(@cloud_cost_analyst_pro)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Exactly. The predictable I/O cost is now a budget line item, not a surprise bill from your cloud provider for memory-optimized instances you don't need.

That local LRU client is key. Without it, you're just paying the I/O tax on every access. With it, you're trading predictable memory overhead per node (for the LRU) for the unpredictable cost of scaling all nodes to hold the full payload.

One metric to track: the hit rate on that local LRU versus calls to the backing store. If it's low, you've misjudged the access pattern and your client is just adding complexity.


cost per transaction is the only metric


   
ReplyQuote
(@calebw)
Reputable Member
Joined: 2 months ago
Posts: 233
 

The "clean, idiomatic way" you're looking for doesn't exist, because the premise is flawed. LangGraph's state dict isn't for your document, it's for your instructions. Treating it like a data bus is how you end up with that hockey stick graph.

Your disk cache with a UUID isn't a step back to 2010, it's admitting the tool's abstraction has a boundary. The real work is in making that client smart. Give it a small in-memory buffer for the chunks the current node is actually working on, and let it manage fetching from the disk backend. The state just holds the UUID and maybe a pointer to the current chunk.

The ugliest part isn't the pattern, it's the boilerplate. You have to write that client, instrument it, and handle cache misses. But that's still cheaper than over-provisioning every node's memory for the once-per-run it actually needs the whole doc.


It's just pattern matching


   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

The boilerplate is the real killer. You end up writing this client that's half state manager, half cache orchestrator, and you have to wire it into every node's init. It feels like you're building a mini-database inside your orchestration layer.

What's worse is when the cache client itself starts holding onto memory because someone wires it up as a global singleton. You see the LRU size set to 100 items, but each item is a 50MB document chunk, and now your "optimization" is just a slower memory leak.

My rule is to make the client stateless per invocation and force it to serialize its pointer map back to the main state. That way a node restart doesn't lose track of what's cached where, and you can actually debug a failed run.


Automate everything. Twice.


   
ReplyQuote
(@integration_ian)
Honorable Member
Joined: 5 months ago
Posts: 396
 

That "cost just shifts" line is the real kicker. It's not just Redis. You're now paying for cache infrastructure, the ops overhead of keeping it running, and the development cost of writing the client glue. All because your orchestration tool decided to abstract away state management, then immediately failed at it.

The tax analogy is perfect. It's a vendor-imposed tax for using their system beyond the happy path. You built your flow with their pretty state dict, now you need a data layer retrofit and the bill comes due.

My team calls this the "integration inheritance tax." You only discover it after you've already committed to the architecture.


Integration is not a project, it's a lifestyle.


   
ReplyQuote
Page 1 / 2