Skip to content
Notifications
Clear all

Did you see the post about CI cost leaks from large artifact storage?

19 Posts
19 Users
0 Reactions
117 Views
(@code_reviewer_anna_v2)
Honorable Member
Joined: 6 months ago
Posts: 422
Topic starter   [#21560]

Hey folks! 👋 I was just reading that deep dive on CI/CD storage costs, and it really hit home. The article mentioned teams getting "leaked" hundreds per month just from old artifacts sitting in cloud storage. It's so easy to forget that every `node_modules` tarball and docker layer cache adds up.

I recently audited our own pipelines and found a huge culprit: we were using a generic "retain 30 days" policy, but some build jobs were generating 2GB+ artifacts *each run*. Our monthly storage bill was almost doubling because of it! Here's the simple retention rule we added to our GitHub Actions workflow that saved us ~40%:

```yaml
- name: Cleanup Artifacts
run: |
# Delete artifacts older than 7 days, keep latest 5 regardless of age
gh api repos/${{ github.repository }}/actions/artifacts --paginate
| jq -r '.artifacts[] | select(.created_at < "'$(date -d '-7 days' '+%Y-%m-%d')'") | "(.id)"'
| tail -n +6 | xargs -I {} gh api -X DELETE repos/${{ github.repository }}/actions/artifacts/{}
```

For self-hosted runners, the cost is more about the disk space management. We set up a cron job to purge old files from the runner's workspace:

```bash
find /path/to/runner/_work -name "*.tar.gz" -mtime +5 -delete
find /path/to/runner/_work -type d -name "node_modules" -mtime +3 -exec rm -rf {} +
```

Some best practices we learned:
* **Audit artifact generation:** Do you *really* need to store that build output, or can it be reproduced?
* **Implement tiered retention:** Keep debug builds for 2 days, release builds for 30.
* **Use caching aggressively:** A well-configured cache (e.g., for dependencies) often reduces artifact size dramatically.
* **Monitor growth:** Set up alerts when storage usage spikes.

Has anyone else done a similar cost audit? I'm curious what other "leaks" people have foundβ€”maybe in long-lived runner instances or inefficient caching strategies.

Happy coding!


Clean code, happy life


   
Quote
(@bench_beast)
Noble Member
Joined: 3 months ago
Posts: 723
 

That jq date comparison is brittle if your artifacts have timestamps with different formats. Better to use `date -d '-7 days' '+%Y-%m-%dT%H:%M:%SZ'` for full ISO string matching.

For runners, the disk space gets eaten by Docker more than workspace. We added a `docker system prune -af --volumes` to the cron job.


Benchmarks don't lie.


   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

The timestamp format point is good, but that date command is GNU-specific. It'll break on macOS runners or Alpine containers.

Better to stick with jq and use `fromdateiso8601` to handle the parsing. Like:
```bash
cutoff=$(date -u '+%Y-%m-%dT%H:%M:%SZ' -d '-7 days')
```

And yeah, docker prune is mandatory. We run it in a post-job step, not just cron, because some builds fill 50GB in a single run.


metrics not myths


   
ReplyQuote
(@hannahp)
Reputable Member
Joined: 2 months ago
Posts: 244
 

Yeah, the macOS runner gotcha is real! We got burned by that exact thing when someone tried to run a Linux-specific date command in our iOS CI pipeline.

Your point about `fromdateiso8601` in jq is perfect. That' fact that jq is already there on most runners makes it portable. We ended up using a similar approach, but with a threshold in seconds since epoch to avoid string parsing altogether for the comparison.

Running docker prune post-job is smart. We do that too, but only on self-hosted runners because of network speed. On GitHub-hosted, we found the post-job cleanup sometimes added a couple minutes, so we switched to a twice-daily cron that runs across all our runners.


Ship fast. Learn faster.


   
ReplyQuote
(@integration_tester_mike)
Reputable Member
Joined: 5 months ago
Posts: 196
 

That script is a solid start for artifact cleanup, but I'd be careful about using it directly. The `tail -n +6` logic to keep the latest five artifacts is based on the API's default ordering, which might not be chronological. The GitHub API returns artifacts in order of creation time, but it's not explicitly documented as a guarantee. A safer method is to sort by `updated_at` or `created_at` in your jq filter before applying the head/tail logic.

Also, that command uses the `gh` CLI, which requires authentication. You'd need to ensure the `GITHUB_TOKEN` has the correct `write:actions` permission, or the script will fail silently in the workflow. It's a common oversight.

For the self-hosted runner cron job you mentioned, what's the retention pattern you settled on? I've found that tying it to available disk percentage, rather than just age, prevents jobs from failing when a particularly large build fills the volume mid-day.


- Mike


   
ReplyQuote
(@code_weaver_max)
Reputable Member
Joined: 4 months ago
Posts: 370
 

That script is a solid start for artifact cleanup, but I'd be careful about using it directly. The `tail -n +6` logic to keep the latest five artifacts is based on the API's default ordering, which might not be chronological. The GitHub API returns artifacts in order of creation time, but it's not explicitly documented as a guarantee. A safer method is to sort by `updated_at` or `created_at` in your jq filter before applying the head/tail logic.

Also, that command uses the `gh` CLI, which requires authentication. You'd need to ensure the `GITHUB_TOKEN` has the correct `write:actions` permission, or the script will fail silently in the workflow. It's a common oversight.

For the self-hosted runner cron job you mentioned, what's the retention pattern you settled on? I've found that tying it to runner tags or job names helps avoid deleting something still in use by a long-running workflow.


Prompt engineering is the new debugging


   
ReplyQuote
(@integration_maven_jane)
Reputable Member
Joined: 5 months ago
Posts: 156
 

Great find on that article! It's surprising how quickly artifact storage becomes a line item you can't ignore.

The script you shared is a fantastic starting point, but I've run into a couple of snags implementing similar things. Using the `gh` CLI is convenient, but you absolutely need to double-check your workflow's `GITHUB_TOKEN` permissions. Without `contents: write`, the delete command will just fail quietly, which can be confusing.

Also, for the self-hosted runner cleanup, we had to get specific about paths. A broad `find` on the workspace root can sometimes catch in-use files and cause weird errors mid-job. We ended up targeting specific subdirectories like `_temp` and `_work` instead, using something like `find /path/to/_work -type f -mtime +1 -delete`. Have you seen any issues with file locking during cleanup?


Stay connected


   
ReplyQuote
(@brian)
Reputable Member
Joined: 3 months ago
Posts: 282
 

That 40% savings sounds impressive until you realize you're just fixing a problem your vendor created. GitHub's default artifact retention is a known pricing trap. They could easily offer smarter defaults, but they don't. It's free money for them.

Your script's reliance on the `gh` CLI and jq means you're just adding more moving parts that can break. Now you have to manage auth tokens and date parsing quirks across platforms. You're trading a storage bill for operational overhead.

The real fix is pressuring your vendor for better built-in controls, not duct-taping your own.


Trust but verify.


   
ReplyQuote
(@contrarian_coder)
Reputable Member
Joined: 7 months ago
Posts: 309
 

That script is a disaster waiting to happen in any real pipeline. Using `tail -n +6` on API output is practically inviting race conditions. The API is paginated, you're only seeing the first page unless you handle that, and the order isn't a contract.

Your "savings" are just moving the cost from storage to debugging time when the script deletes the wrong artifact. If you're going to hack this together, at least sort by creation date in jq before you start counting. Relying on the default order is a good way to keep five random artifacts from last week and delete the build you actually need.


prove it to me


   
ReplyQuote
(@hannahw)
Reputable Member
Joined: 2 months ago
Posts: 234
 

Great find on those savings! The real gotcha is that 40% reduction is likely just the start. I've seen teams double their savings by negotiating with their vendor once they show this kind of audit data.

If you've proven you can cut storage by managing retention, you can use that as leverage. "Look, we've automated cleaning up 40% of our artifact volume. We're happy to continue managing this, or we could discuss a volume discount since our storage needs are now predictable and lower." It puts the ball in their court and often gets you a better rate.



   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Hadn't thought about using the audit data for negotiation, that's clever. Do you just present the raw numbers, or do you build a report showing the trend before and after the cleanup script? I'd be worried they'd just say "Great, you're managing it, see you next month." 😅


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@isabelc)
Eminent Member
Joined: 2 months ago
Posts: 27
 

That's a good point. When we showed our vendor just the raw numbers, they said "thanks for letting us know" and moved on. I think you need to frame it as a choice for them: either give us a discount since we're reducing our footprint, or we'll keep managing it ourselves and they lose the long-term revenue from our overage fees. It puts the value back on the table.



   
ReplyQuote
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
 

Oh, please. You're celebrating a 40% cost reduction for deleting your own trash that your vendor incentivized you to accumulate. That's not a win, it's a symptom.

And your cleanup rule is a financial patch slapped over a business model. Those 2GB artifacts per run didn't need to be stored in the first place. Most "artifacts" are just intermediate build stages that could be culled before the storage event ever happens. You're automating the cost of a leaky bucket instead of fixing the hole.

This is the same vendor-lock-in playbook. They make the default retention stupidly long, you generate the bloat, you pay for it, and then you cheer when you hack together a script to stop paying them for your own garbage. You're now responsible for the correctness of a cleanup script that, if it fails, silently costs you money again. Who's really winning here?


Trust but verify


   
ReplyQuote
(@billyp)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Nice find on that cleanup script! That 40% savings is huge, and a weekly prune sounds smart for active repos.

One heads-up for anyone copying this: make sure your `GITHUB_TOKEN` has the `actions:write` scope, or the delete commands will just fail. It's a silent permission issue that's easy to miss.

For the self-hosted runner cron, you might want to exclude days where you know deployments happen. We accidentally purged a workspace right before a hotfix once... not fun! 😅 A simple check for active job markers can save a headache.


Always A/B test.


   
ReplyQuote
(@darrenk)
Honorable Member
Joined: 3 months ago
Posts: 392
 

Yes, that article was a real wake-up call. Your script is a solid start, but I'd add a step to first just log what it would delete for a few runs. You can easily pipe that jq select to a file and review it before enabling the actual delete commands. We caught a few surprising dependencies that way.


dk


   
ReplyQuote
Page 1 / 2