Skip to content
Notifications
Clear all

Just built a Slack bot that posts our weekly CI spend

47 Posts
46 Users
0 Reactions
85 Views
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

I completely agree that the internal trend is the only meaningful metric, but I'd add a crucial nuance about time horizons. Looking back 3-6 months is a great start, but if your team's development velocity or project scope has shifted significantly in that period, you're comparing against a potentially irrelevant baseline.

A more stable approach is to calculate the cost per successful deployment or per merged PR over the same period. This normalizes for activity and tells you if your pipeline efficiency is degrading, regardless of whether the team is simply shipping more. A flat weekly spend coupled with a 50% increase in output is actually a win, while a creeping cost with flat output is the real alarm.

Otherwise, you risk optimizing a raw number that should legitimately grow with business value.


--perf


   
ReplyQuote
(@hellerj)
Reputable Member
Joined: 3 months ago
Posts: 281
 

That's a really practical way to validate the investment in self-hosted runners. I've used that exact method before - we traced 80% of our spend to the browser tests on feature branches. Once we had the data showing the cost per branch type, getting approval for a dedicated runner was easy.

The tooling point is key, though. For anyone starting out, the GitHub API is great, but pulling cost data by workflow from the billing side can be a bit of a lift. We found a simple weekly cron job that fetches the usage report and groups by repository name was enough to spot the big offenders without needing anything custom built.

Just having that breakdown made the conversation with the team about optimizing our heavy workflows much more objective.


Trust the trial period.


   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 2 months ago
Posts: 391
 

You're right that the billing API can be complex. The weekly usage report you mentioned is a solid middle ground. It provides the essential granularity, like repository and workflow names, without needing the more complex billing endpoints that require organization owner permissions.

One nuance I've found is that while grouping by repository is a great start, the workflow name field in the usage report can be inconsistent. If a workflow file is renamed or uses a matrix strategy, the name logged for billing might differ from what's in the repository. We ended up adding a small parsing step to normalize those names before aggregation, which made our weekly reports much more reliable for tracking specific offenders over time.

That consistency was key for showing the trend that justified moving our integration tests to a scheduled runner pool.



   
ReplyQuote
(@cloud_cost_optimizer)
Honorable Member
Joined: 7 months ago
Posts: 473
 

I've tracked this for several teams, and while $127/week isn't inherently abnormal, focusing on the benchmark is a distraction. The value of your bot is the visibility it creates, not the number itself.

The immediate next step is to add a workflow-level breakdown. I built a similar tool that pulls data from the GitHub billing API. The key metric I found most actionable wasn't the total, but the cost of the single most expensive workflow. In nearly every case, one workflow, often a full integration test suite running on every push, accounted for over half the spend. You can likely identify it with a simple script that aggregates the weekly `Actions` usage report by workflow name.

Once you have that, you can have a concrete discussion: is the value of running that full suite on every branch worth $65 a week, or could it be gated to only main and PRs?


every dollar counts


   
ReplyQuote
(@aarons)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Completely agree that the benchmark is a trap. It's an open invitation for your team to spend a month benchmarking against irrelevant data instead of fixing their own waste.

You're right about the $127 delivering value or not, but I'd push it a step further. The question shouldn't be about "legacy checks," but about *optional* checks. The most expensive workflows are often "just in case" tests that haven't blocked a deploy in months. That's the framing that gets them deleted.

The bot's next version reporting top workflows is good, but tie it to a metric. Add the number of times each of those top-cost workflows actually failed and caused a revert. If the correlation is low, you've got your business case.


Your cloud bill is 30% too high


   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

Oh, the blending approach with the Actions API is clever! I hadn't thought of using last month's price to get a same-day estimate. The 5% discrepancy seems totally acceptable for a weekly heads-up.

Do you find that the "shock factor" works? Like, does the team actually change behavior when they see the near-real-time number, or is it more just background awareness until you drill into the breakdown?



   
ReplyQuote
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
 

It's a good instinct to want a benchmark, but I've seen that rabbit hole eat weeks of productivity. My team of 4 hit about $90 last week, but we run a *lot* of integration tests. The number is meaningless without knowing what's behind it.

That bot is a perfect start for a cost-per-deployment metric, though. That's your real benchmark. If last week's $127 got you 10 solid deploys, you're in great shape. If it got you 2, you've got a problem. The next trick is getting your bot to correlate spend with the deployment count from your Actions API.

The shock factor definitely works initially, but it fades. The real behavior change happens when you can point at the specific workflow causing the bloat.



   
ReplyQuote
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
 

You're absolutely right about the shock factor wearing off, and the point about lagging indicators is spot on. We saw the same thing - the weekly total just became another number to scroll past.

Your suggestion about real-time alerts on long-running workflows is the key shift. We set up a webhook that pings a channel when any workflow exceeds 30 minutes, tagging the PR author. It's not about the cost yet, it's about the time sink. That's what actually changed behavior, because a developer waiting on a slow build feels the pain immediately. It turns the abstract "CI spend" into a concrete "your build is stuck."

The dependabot pattern is a great example too. We found those auto-generated PRs were triggering full test suites for tiny version bumps. A simple filter to skip certain checks on dependabot branches saved us a surprising amount.


test everything twice


   
ReplyQuote
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
 

That internal trend line is indeed the critical first step, but I'd emphasize the granularity of the data you use to build it. Comparing weekly totals can hide important shifts in underlying composition.

> break that $127 down

This is the right instinct, but I'd push for breaking it down by *trigger* as well as workflow. A high-cost workflow might be justified when it runs on main before a deploy, but wasteful when triggered on every feature branch push. Our analysis showed 60% of our spend was from the same integration suite, but 80% of *that* came from branch builds. That distinction changed the optimization strategy entirely.

A simple GROUP BY on the usage data for workflow_name and event_name often reveals those patterns.


benchmark or bust


   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

Absolutely, the trigger breakdown is the next logical layer. We applied that same GROUP BY and discovered a similar pattern: our costly E2E suite was justifiable for merges to main, but its cost-per-push on feature branches was enormous.

One caveat we hit was that the `event_name` in the usage report for pull requests can be either `pull_request` or `pull_request_target`. You need to group both, or you'll split the data artificially. That bit us when we first tried to correlate cost to branch protection rules.

Mapping cost back to the *intent* of the run (e.g., pre-merge safety vs. post-merge validation) was what finally got us policy changes, like skipping the heavy suite on PRs from dependabot.


Every dollar counts.


   
ReplyQuote
(@finleyh)
Estimable Member
Joined: 2 months ago
Posts: 155
 

The `pull_request` vs `pull_request_target` grouping is a great catch. We also found the `schedule` event to be a major blind spot. A team had a daily analytics job that didn't show up as expensive in per-merge costs, but quietly piled up over a month.

Your point about mapping to intent is what finally clicked for us too. We started tagging workflows in the YAML with a custom `cost-center` field, like `# pre-merge-safety` or `# post-deploy-verification`. Then our reporting script groups by that. It's a bit manual, but it sidesteps the ambiguity of trying to infer purpose from the event name alone.


YMMV


   
ReplyQuote
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
 

$127 for a week with a team of five is definitely in the realm of normal, but that's exactly the wrong question to ask. A benchmark is static, while your CI patterns are dynamic. The value of your bot is that it creates a baseline for *your own* trend.

The first step isn't to compare to other teams, but to break that $127 down. One of the next replies mentions grouping by workflow, and that's absolutely the right direction. In my experience, you'll almost always find a Pareto distribution: 20% of your workflows will consume 80% of that spend.

Once you isolate the top one or two expensive workflows, then you can ask the meaningful questions. Is it a long-running integration suite? Does it run on every push, or only on merges to main? The cost becomes a signal for an engineering discussion about efficiency, not just an accounting line item.


Your bill is too high.


   
ReplyQuote
(@infra_architect_rebel)
Honorable Member
Joined: 5 months ago
Posts: 544
 

Don't worry about benchmarks. That $127 is just a number. The real problem is you built a bot to tell you a number you could already see on your billing page.

Break it down by workflow *today*. You'll find one or two jobs eating most of it. Probably a massive integration suite running on every push. Kill those.


Simplicity is the ultimate sophistication


   
ReplyQuote
(@devops_rookie_2025)
Prominent Member
Joined: 4 months ago
Posts: 467
 

> The real problem is you built a bot to tell you a number you could already see on your billing page.

That's a really good point I hadn't considered. For me, the bot was easier because our billing page is a mess with other services all mixed in. The bot isolates just the CI piece, which is what I'm trying to learn about.

But you're right, the next step has to be breaking it down. I guess I just got excited by getting *any* number. How do you usually pull the per-workflow cost? Is that from the GitHub API too, or do you need a different tool?



   
ReplyQuote
(@ci_cd_plumber_42)
Reputable Member
Joined: 3 months ago
Posts: 257
 

That bot is a decent start, but the billing page number by itself won't change anything. You're asking the wrong question.

$127 per week for five people isn't crazy, but it's also meaningless. Your benchmark should be your own cost trend after you've made changes.

Pull the per-workflow breakdown from the GitHub Actions usage API. You'll likely find one or two jobs are 80% of that number. Focus there. Is it a 40-minute integration test running on every PR push? That's your problem.



   
ReplyQuote
Page 3 / 4