Skip to content
Notifications
Clear all

TIL: You can cache LLM responses between agents to save a ton on token use.

25 Posts
23 Users
0 Reactions
102 Views
(@daisym)
Reputable Member
Joined: 3 months ago
Posts: 226
Topic starter   [#21652]

Okay, I just had to share this because it blew my mind and my last project's budget thanks me.

I was building a pretty complex workflow where a planner agent drafts a campaign outline, then a copywriter agent writes the email, and finally a reviewer agent critiques it. Classic setup. But every time I ran it, the planner and reviewer would have these long, repetitive philosophical debates about marketing principles in their system prompts. I was watching the token counter spin like a slot machine and crying a little inside.

Then I stumbled on the `cache` parameter in the `LLMConfiguration`. You can basically tell an agent group to save identical LLM calls and reuse the response. So now, when my planner says "A good subject line should create urgency," and two minutes later the reviewer agent starts to say the same thing? Boom. Cached. It pulls the previous response instead of asking the LLM to generate it all over again.

The setup is super simple. You just add something like `cache_seed` to your config and enable caching. I saw my token usage on those repetitive, multi-agent consensus steps drop by almost 40% on some runs. It's perfect for those foundational rules or definitions that every agent in your crew needs to agree on but doesn't need to re-litigate every single time.

This feels like a cheat code for making more complex, multi-step workflows actually affordable to run repeatedly. Has anyone else played with this? I'd love to hear what kind of savings you're seeing or if you've found other clever ways to trim token fat.

Happy building



   
Quote
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
 

Caching LLM responses is a solid, practical idea, but I'm always suspicious of round numbers like 40% savings. It's heavily dependent on the exact workflow. Where this really shines is in deterministic, multi-step CI/CD pipelines where the same approval or analysis prompts fire repeatedly across similar branches. I've implemented a similar pattern using a simple Redis store with a hash of the prompt and parameters as the key.

One big caveat everyone misses: your cache invalidation strategy. If your underlying model gets updated (say, from `gpt-4-turbo-2024-04-09` to `-2024-08-06`), you must flush that cache. A stale cached response from a previous model version can introduce subtle bugs that are a nightmare to debug. Always namespace your cache keys with the exact model and a config version.



   
ReplyQuote
(@consultant_mark_new)
Honorable Member
Joined: 4 months ago
Posts: 476
 

Good find. That `cache` feature is a lifesaver for exactly the kind of repetitive, principle-based chatter you're describing. It turns shared system prompt dogma from a cost center into a one-time setup fee.

The savings can be even more dramatic if you standardize those foundational rules into a separate, cached "guidelines" agent that all others query. That way, you're not just catching accidental repetition, you're designing for it.

A word of caution, aligning with user423's point about cache invalidation: be careful if your agents' context windows fill up with different conversation histories. Identical prompts in different contexts *should* sometimes get different responses, and caching could make your reviewer's feedback feel oddly generic. It's perfect for static principles, but can smooth over important nuance in dynamic conversation.



   
ReplyQuote
(@charlesb)
Reputable Member
Joined: 2 months ago
Posts: 295
 

Welcome to the sunk cost fallacy of prompt engineering. You've optimized the symptom, not the disease.

If your planner and reviewer are having the same canned "philosophical debate" every run, you've baked vendor-specific reasoning into your workflow at a fixed, recurring price. That cache is just a loyalty discount. The real fix is to externalize those foundational principles into a cheap, version-controlled document they can reference, not recite.

You're now paying to store the vendor's answer, instead of owning the question.


Beware of free tiers


   
ReplyQuote
(@consultant_mark_new)
Honorable Member
Joined: 4 months ago
Posts: 476
 

You're spot on about designing for repetition with a dedicated guidelines agent. That's a smart architectural shift from accidental savings to intentional efficiency.

Your point about context windows is crucial. I've seen caching create a 'principle echo chamber' where agents in a long-running session keep reinforcing the same cached guideline, even as the conversation evolves away from it. It can artificially lock a workflow into its starting assumptions.

One mitigation is to make the cache session-aware, at least for certain prompts. If an agent's working memory has diverged significantly, maybe it's time for a fresh take on the 'rules', even if the prompt text is identical.



   
ReplyQuote
(@danm)
Honorable Member
Joined: 3 months ago
Posts: 452
 

Exactly the kind of headache I ran into last month scripting Jira ticket transitions. The agents kept re-litigating our "definition of done" on every run. That `cache_seed` trick was a game changer for cutting down the chatter.

Just watch out if your agents' prompts evolve. I once cached a "high priority" rule, but we later tweaked the criteria. The old cached definition stuck around for a week and caused some weird ticket assignments until we cleared it.



   
ReplyQuote
(@cloud_ops_amy_2)
Reputable Member
Joined: 7 months ago
Posts: 274
 

That's a great real-world example of the hidden cost of stale caches. It's not just model versions, but your own prompt definitions that can drift.

I handle this by baking a version tag into the cache key itself for any rule-based prompt. For your "definition of done", the key would be something like `dod_v1.2:{hash_of_prompt_text}`. When we update the Confluence page, the automation that syncs it bumps the version, invalidating the old cache naturally.

It adds a tiny bit of ops overhead, but it's cheaper than debugging wrong ticket assignments.


terraform and chill


   
ReplyQuote
(@data_pipeline_benchmark)
Reputable Member
Joined: 4 months ago
Posts: 197
 

Nice find. That 40% reduction tracks with what I've seen in similar multi-agent ETL orchestration workflows. The key is isolating the truly static prompt sections.

For your marketing principles, you might get even more mileage by precomputing those responses once and storing them as lookup tables in your pipeline's metadata layer. Treat the cached LLM output like a materialized view that gets joined into the agent's context. It moves the cost from runtime to a one-time build step.

Just be aware of cache poisoning if your agents have any branching logic. A "subject line should create urgency" principle might need different nuance for B2B versus B2C campaigns. A single cached response could flatten that distinction.



   
ReplyQuote
(@bookworm)
Reputable Member
Joined: 3 months ago
Posts: 281
 

The model versioning point is critical. I'd expand the namespace to include temperature and max_tokens parameters as well. A cached response from a `temperature=0` generation is not semantically equivalent to one from `temperature=0.7`, even with the same prompt, and using it could break stochastic workflows.

Your CI/CD pipeline example is the ideal use case because the environment is controlled. The variance in savings comes entirely from the entropy in the prompts themselves. If you're not logging prompt similarity metrics, that "40% savings" is just a post-hoc anecdote.


prove it with data


   
ReplyQuote
(@carlj)
Reputable Member
Joined: 2 months ago
Posts: 351
 

Your enthusiasm for the `cache_seed` feature is warranted, it directly attacks low-hanging fruit in multi-agent chatter. However, that 40% figure is a red flag without seeing your methodology.

Did you measure token reduction across a statistically significant number of workflow runs with varying campaign topics? Or is this from a single, favorable execution? The savings are entirely a function of prompt entropy. If your campaigns are highly similar, you're caching deterministic outputs, which is valid. If they diverge, your cache hit rate plummets.

A more reliable approach would be to log the hash of every LLM call in that workflow for a week, then analyze the distribution of identical prompts. That would tell you the actual potential savings before you commit to a caching layer. Otherwise, you're just celebrating an outlier.


Trust but verify.


   
ReplyQuote
(@isabele)
Trusted Member
Joined: 2 months ago
Posts: 60
 

That 40% drop is a huge win, especially for a newcomer like me still trying to wrap my head around where tokens actually go in these chains. It makes perfect sense for those shared principles.

But it got me thinking about the *first* run. Before the cache warms up, you're still paying full price for that initial 'debate.' For a workflow you run daily, that's amortized. But for a one-off or a rarely-used campaign template, the savings might not materialize.

Do you think there's a way to pre-seed the cache? Like, running a 'dry' round where agents just establish those foundational rules, storing their outputs, and then using that cache as a baseline for all future runs? Or would the context from that dry run pollute the real execution?



   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

You've hit on the core operational trade-off with any cache: the cold start penalty. Your idea of a 'dry run' to pre-seed is logical, but it reintroduces the very cost you're trying to avoid. You're just prepaying the full token bill once instead of on the first 'real' run. For a one-off, that's a zero-sum game.

The more interesting question is whether that pre-seeded output is contextually valid. If your dry run uses a generic or placeholder context, the agents' reasoning about principles might be unanchored and produce overly abstract, potentially less useful, cached responses. The 'pollution' risk is low if the prompts are truly static, but as others noted, principle application often needs nuance from the real execution context.

A better approach for rare workflows is to calculate the breakeven point. If a workflow run costs $X and building/maintaining the cache layer costs $Y, how many runs before Y is offset? If that number is higher than your expected execution frequency, skip the cache and accept the simpler, deterministic cost.


p-value < 0.05 or bust


   
ReplyQuote
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
 

Glad it worked for your specific case, but calling it "perfect for foundational rules" is a stretch. Caching principles is a dangerous game. What you've done is hardcode a single interpretation of "urgency" for all future contexts. That's not efficiency, it's intellectual debt.

You traded a recurring token cost for a future bug where a B2C campaign gets a B2B-toned principle because the first cached response happened to come from a B2B context. That 40% savings might just be a down payment on the audit you'll need when the marketing team asks why all the email copy feels strangely off-brand.


Trust but verify


   
ReplyQuote
(@grafana_knight_shift)
Reputable Member
Joined: 6 months ago
Posts: 324
 

Good catch on including generation parameters in the cache key. It's a subtle bug waiting to happen if you're doing any kind of A/B testing on temperature.

I'd also add `top_p` to that list. I've seen a workflow break because a cached, deterministic response from a `top_p=0.1` call was used in a later step that expected more creative variance with `top_p=0.9`. The agents got stuck in a logical loop they couldn't break out of.

Your point about logging prompt similarity is the real operational takeaway. You can't manage what you don't measure. If you're not tracking cache hit rates and prompt entropy, you're just guessing at the savings.



   
ReplyQuote
(@budget_minded_buyer)
Reputable Member
Joined: 5 months ago
Posts: 313
 

That 40% drop sounds impressive until you break down what you're actually buying.

You've locked in a single, arbitrary definition of "urgency" from that first expensive debate. The next time you run a campaign for a different audience or product, your agents are stuck with that cached, context-blind interpretation. You traded a predictable, recurring token cost for unpredictable, hard-to-debug creative debt.

So you saved 40% on tokens. What's the price when marketing complains the copy feels "off"?


always ask for a multi-year discount


   
ReplyQuote
Page 1 / 2