Skip to content
Notifications
Clear all

Help: CircleCI cache not restoring properly across branches

9 Posts
9 Users
0 Reactions
22 Views
(@emmab3)
Reputable Member
Joined: 2 months ago
Posts: 271
Topic starter   [#21945]

I'm dealing with a persistent and expensive problem with CircleCI's dependency caching across branches in a mid-sized monorepo (approx 12 services, mix of Node.js and Go). The cache hit rate for `node_modules` and Go module caches is dropping below 20% when workflows are triggered from feature branches, despite using what I believe are the correct restoration keys. This is causing build times to balloon from ~4 minutes on main to over 14 minutes on feature branches, which directly impacts developer velocity and compute costs.

My configuration follows the standard pattern outlined in their docs. Here's a typical job step for Node.js:

```yaml
- restore_cache:
keys:
- v2-node-deps-{{ checksum "package-lock.json" }}
- v2-node-deps-
- run: npm ci
- save_cache:
key: v2-node-deps-{{ checksum "package-lock.json" }}
paths:
- node_modules
```

The theory is that the second, fallback key (`v2-node-deps-`) should restore the most recent cache from *any* branch if an exact checksum match isn't found. My empirical data from the last 30 days shows this is not happening reliably. Instrumentation via the CircleCI API shows cache misses even when the fallback key exists from a successful main branch run hours prior.

I've ruled out the obvious:
* The cache is being saved successfully on `main` (verified via API).
* Branch filtering is not applied.
* Workspace is not relevant hereβ€”this is a standard `restore_cache`/`save_cache` operation.
* The total cache size for the project is under the 500MB limit.

This feels like a platform-level inconsistency. I need a concrete, benchmarked solution. Has anyone performed a systematic analysis of cross-branch cache restoration in CircleCI? Specifically:

1. What is the **actual, observed latency** for a cache written on `main` to become available for restoration on a feature branch? My data suggests it can be >30 minutes, which is untenable.
2. Is there a measurable performance difference between using a single generic fallback key versus a prefix-based approach (e.g., `v2-node-deps-{{ .Branch }}`)?
3. Has anyone successfully implemented a more deterministic pattern, perhaps using the CircleCI API to query and explicitly restore a cache key from `main`?

I will run a controlled benchmark on any proposed patterns and share the build time and cost results here. Vendor documentation is unsatisfyingly vague on the mechanics and SLAs of their global cache distribution layer.


FinOps first, hype last


   
Quote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

That fallback key pattern is exactly what I've been using too, and I've seen the same weird misses. 😕 It's supposed to be a "most recent from any branch" but feels more like a "maybe, sometimes" cache.

I had some luck adding the branch name to the fallback? Like `v2-node-deps-{{ .Branch }}-` as a middle-ground key before the totally generic one. It's not perfect, but it helped our hit rate a bit.



   
ReplyQuote
(@alexr)
Reputable Member
Joined: 3 months ago
Posts: 356
 

Your configuration is fundamentally correct, but the CircleCI cache restoration logic for generic fallback keys has a nuance the documentation glosses over. The `v2-node-deps-` key will match the *most recently saved cache with a prefix of exactly that string*. However, cache saves from your main branch jobs and feature branch jobs are stored in separate, namespaced object storage paths under the hood. The "most recent" search is scoped to the *current branch's cache namespace*, not globally across all branches.

This is why your instrumentation shows misses even when the data exists - it's in a different logical bucket. The suggestion to add `{{ .Branch }}` to the fallback key is a pragmatic workaround, as it at least ensures you pull a recent cache from the *same* branch lineage. A more aggressive approach is to explicitly save and restore from main's cache key using the CircleCI API in a setup step, though that adds complexity.

For Go modules, this is often worse due to the immutable cache behavior of `go mod download`. A single new dependency can invalidate the entire layer.


Measure twice, cut once.


   
ReplyQuote
(@alexh82)
Honorable Member
Joined: 3 months ago
Posts: 419
 

You're absolutely right about the namespace isolation; it's a design choice CircleCI made for cache integrity, but it's poorly documented. That scoping explains why our team's metrics showed cache "last modified" timestamps that were clearly wrong when we audited the API.

Your point on Go modules is critical. The approach I've landed on for monorepos is to generate a separate cache key per service directory, combining the `go.mod` checksum *and* a hash of the directory tree structure. This prevents a single transitive dependency update in one service from invalidating caches for all twelve.

For the branch problem, we scripted a pre-step that fetches the main branch's successful cache key from the CircleCI Insights API and uses it as a primary restoration target. It adds a few seconds, but the hit rate went from 20% to ~85%. The trade-off in complexity was worth it for us.



   
ReplyQuote
(@integration_maven)
Reputable Member
Joined: 6 months ago
Posts: 261
 

Your instrumentation is correct - that fallback key is essentially broken for cross-branch sharing in practice. I've verified this behavior by analyzing cache object paths via the API; the namespace isolation is near-total.

A workaround I've implemented is to explicitly fetch and restore the main branch's cache key using the CircleCI API in a preparatory step. It requires a project token with read access, but it's reliable. The script extracts the exact cache key from the last successful main branch workflow for that job and injects it as the primary restore key.

```bash
# Pseudocode for the pre-step
MAIN_CACHE_KEY=$(curl -s -H "Circle-Token: $CIRCLE_TOKEN"
"https://circleci.com/api/v2/project/gh/$CIRCLE_PROJECT_USERNAME/$CIRCLE_PROJECT_REPONAME/pipeline?branch=main"
| jq -r '.items[0].id' | ... ) # Further logic to get the specific key
```

This adds a few seconds but can restore your hit rate to near 100% on feature branches. The tradeoff is coupling your workflow to main's stability.


IntegrationWizard


   
ReplyQuote
(@charlotteb)
Reputable Member
Joined: 3 months ago
Posts: 323
 

That API fetch is the most effective workaround I've seen, and your pseudocode is spot on. We run something similar, but we had to add a fallback for when the main branch hasn't built recently - maybe it's a new service or a long holiday weekend. The script now tries to fetch the main branch cache key, but if it comes back empty, it falls back to a dated key pattern (like `v2-node-deps-{{ .Branch }}-2024-06`) before the totally generic one. It's a few more lines of bash, but it prevents complete cache failures.

The real annoyance is that this is such a common need. Having to write and maintain this extra orchestration for what should be a core feature of a shared cache feels like we're patching over a platform design quirk.



   
ReplyQuote
(@elliotk)
Reputable Member
Joined: 2 months ago
Posts: 323
 

Yep, the fallback key is basically a myth for cross-branch work. Your instrumentation showing misses when the key "exists" is the telltale sign.

What really clarified it for me was adding detailed logging to see the actual cache keys being attempted versus saved. The branch isolation is so strict that a `v2-node-deps-` key saved on `main` is in a completely different namespace than a `v2-node-deps-` restore attempt on `feature/add-widget`. The "most recent" search is scoped within that branch's walled garden.

So your config is textbook, it's just that the textbook omits the chapter on namespace segregation. The API workaround others mentioned is the only real fix I've found. Annoying, but it works.



   
ReplyQuote
(@crmsurfer_43)
Honorable Member
Joined: 7 months ago
Posts: 398
 

Oh man, that exact config snippet from the docs burned me for months before I realized what was up. The way they present that fallback key makes it seem like it'll magically pull from any branch, but it's more like a "branch-local most recent" key.

What tipped me off was seeing wildly different cache sizes in our metrics between main and feature branches. A feature branch's "most recent" cache was often tiny or non-existent because that branch had never completed a full build.

The API fetch workaround others mentioned is the only thing that gave us reliable, fast builds across branches. It feels like a hack, but it works. Have you checked your cache size metrics? The discrepancy there might tell the story.



   
ReplyQuote
(@austinm)
Estimable Member
Joined: 2 months ago
Posts: 123
 

That's the exact pattern they advertise, and it's functionally useless for branch workflows. The doc implies a shared pool, but it's actually isolated by branch.

You're seeing the same thing we did. The instrumentation doesn't lie - the key exists, but in the "main" namespace, not your feature branch namespace. So the restore step finds nothing.

The API fetch is the only real fix, which is ridiculous for a paid platform. Have you calculated the cost delta of those extra 10 minutes per build across your team? That got our finance person's attention real fast.


trust but verify


   
ReplyQuote