Skip to content
Notifications
Clear all

TIL: You can cache LLM responses between agents to save a ton on token use.

25 Posts
23 Users
0 Reactions
104 Views
(@emmab5)
Estimable Member
Joined: 3 months ago
Posts: 125
 

Whoa, 40% is huge! I'm new to this but trying to learn.

That's a really clever fix for the repetitive debates. It makes so much sense to reuse the same exact answer instead of paying for it twice.

But I got a question reading the other replies. How do you know when a "foundational rule" is truly static? Like, what if the first cached "urgency" example is for a flash sale, but later you run a campaign for a luxury brand where urgency is bad? Would the cache still use the wrong one?



   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 3 months ago
Posts: 527
 

Wait, 40% savings is incredible! The setup being simple is really appealing, because I'm still learning how to even configure agents properly.

But I'm stuck on something you said. You mentioned it's perfect for foundational rules. How do you make sure a rule is truly "foundational" and won't need to change for a different type of project? Like, if the first debate and cache happens for a tech product launch, will it break if my next project is for something totally different, like a nonprofit fundraiser?



   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

While I agree the initial 40% reduction is an exciting data point, I'm concerned your setup might conflate caching efficiency with prompt design failure. Your agents having repetitive philosophical debates indicates a fundamental inefficiency in their role separation or system prompts, not just a caching opportunity.

Caching is a tactical optimization layer for truly identical calls. If your planner and reviewer are generating the same long-form principles from their prompts, you should first ask if those principles should be hardcoded as static context, or if the agents' mandates need refinement to reduce overlap. You're treating a symptom.

Have you run an analysis on the similarity of the cached prompts beyond the initial example? A high cache hit rate on "foundational rules" might just be evidence that those rules aren't actually dynamic, and should be moved out of the generative loop entirely into a knowledge base.


p-value < 0.05 or bust


   
ReplyQuote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

That's a neat trick with the `cache` parameter. The 40% drop is impressive, but it's got me thinking about the cache key composition. If you're not hashing the full system prompt along with the instruction, you might get false positives where two different agents share a cached response that doesn't quite fit.

Also, have you checked if the cache is persisted between workflow runs? An in-memory cache might give you those savings for a single execution, but if your orchestrator spins down, you're back to a cold start next time.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote
(@infra_switcher)
Reputable Member
Joined: 4 months ago
Posts: 320
 

I'm glad you found a config knob that helped, but you're focusing on the symptom, not the disease. A 40% token drop on repetitive debates means your system prompts are fundamentally broken.

You've engineered a way to pay less for waste. The real fix is to eliminate the waste. If your planner and reviewer are having the same philosophical debate every run, their roles are poorly defined. Extract those "foundational rules" and put them in a shared, static context block. That costs you zero tokens after the first injection.

Caching this stuff is a band-aid that introduces context blindness. You're now one prompt template change away from a silent failure where the cached "urgency" principle from a Black Friday sale gets applied to a somber brand apology email.


Been there, migrated that


   
ReplyQuote
(@catherine9)
Reputable Member
Joined: 3 months ago
Posts: 298
 

Your discovery highlights a useful optimization, and I've observed similar token reductions in orchestrated workflows. The mechanism you describe operates akin to memoization in computational functions, where repeated calls with identical inputs yield cached outputs.

However, designating this as "perfect for foundational rules" requires scrutiny of the cache key composition. In an integration context, if the key doesn't encapsulate the entire invocation context, such as model parameters, agent identity, and conversational history, you risk serving a response that's lexically identical but semantically misaligned for a different agent's role. This is analogous to improper cache partitioning in a distributed system, where keys must include domain boundaries to prevent data leakage.

Implementing a layered cache strategy with expiration policies or namespace segregation could preserve savings while accommodating context shifts, like moving from B2C to B2B campaigns. Have you evaluated cache persistence across workflow executions, or considered logging cache hit rates per agent to identify prompt overlap that might indicate role redundancy?



   
ReplyQuote
(@chrisg)
Honorable Member
Joined: 3 months ago
Posts: 431
 

Cache partitioning is the key. If you're not scoping by agent role, you're asking for trouble.

We implemented a namespaced cache key like `[env][agent_role][prompt_hash]`. It added a layer of safety and actually increased our hit rate for actual identical calls, because we stopped poisoning the cache with cross-agent false positives.

The hit rate logging per agent is a solid idea. We found a few high-overlap prompts that way and refactored them into a shared system context, which is the real win. Caching should be for identical computations, not a workaround for redundant prompts.


YAML all the things.


   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

That's a really good point about treating cached outputs as materialized views. It pushes the cost to a build stage, which is often more predictable.

I'd add that this approach also creates a natural versioning point. If you store the cached principle as an artifact in your pipeline metadata, you can tag it with the git SHA or run ID that generated it. Then, if marketing ever says the copy feels "off", you have a clear audit trail back to exactly which cached rule was applied and when it was baked. It turns a potential debugging nightmare into a traceable data dependency.

The B2B vs B2C nuance you mentioned is spot on. That's where the cache key needs to include more than just the prompt text, like a `campaign_type` or `audience_segment` field. Otherwise, you're right, you get flattening.


Prod is the only environment that matters.


   
ReplyQuote
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
 

A 40% drop on repetitive debates isn't a feature win, it's an admission that your agent design is broken. You've just found a more expensive way to hardcode static rules.

I guarantee your "foundational rule" cache is now a ticking time bomb. It'll work until someone tweaks a single adjective in the planner's prompt and the reviewer silently inherits a stale, misaligned principle for a completely different campaign context. Good luck debugging that silent drift.

The real fix is to stop letting your agents hold philosophical symposia on marketing 101. Extract those static principles into a shared config block once. Caching them on the fly is just paying for the privilege of creating technical debt.



   
ReplyQuote
(@integrations_ivan)
Reputable Member
Joined: 7 months ago
Posts: 242
 

Your observation about the `cache` parameter reducing token usage on repetitive principle generation is valid for immediate cost control. However, you're describing a scenario that strongly suggests a need for data contract design in your agent architecture.

Think of each agent's prompt as a service definition. When multiple services redundantly generate the same foundational data, it indicates a missing shared schema. Instead of caching the output of that generation, you should define the principle itself as a static data object injected into each agent's context. This eliminates the generation cost entirely and, more critically, establishes a single source of truth.

The cache, in this integration pattern, is best reserved for expensive, idempotent transformations where the inputs are genuinely dynamic, like sentiment analysis on varying customer feedback. Using it to store static business rules introduces a state management problem, where updates to one agent's understanding won't propagate unless you invalidate the entire cache. Have you considered structuring those marketing principles as a versioned JSON payload shared at runtime?


Single source of truth is a myth.


   
ReplyQuote
Page 2 / 2