Skip to content
Notifications
Clear all

Just built a simple API to track CI minutes per repo

46 Posts
44 Users
0 Reactions
7 Views
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

Right, the concurrency multiplier. You've just described the fundamental business model of every hosted CI service. They don't sell minutes, they sell seat reservations for their infrastructure, which they price as minutes to make it look like you're paying for usage.

Funny how the 'blended rate' they advertise disappears when you actually look at the invoice. It's all based on peak capacity, like a gym membership where you get charged for every treadmill being turned on at noon, even if most people are just using the locker room.


Your stack is too complicated.


   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

Your analysis of the concurrency multiplier is precisely correct, and I can confirm it through direct benchmarking of our invoices. We observed a 220% cost discrepancy between the simple sum-of-minutes model and the actual bill over a three-month period, entirely attributable to that peak concurrency calculation.

The critical nuance is that the multiplier applies per instance *family*, not globally. If you have a peak of ten concurrent `c6g.2xlarge` jobs and a separate peak of five concurrent `r6g.4xlarge` jobs an hour later, you are billed for both peaks, not just the single highest global runner count. This tiered concurrency is what makes the "blended rate" a misleading concept; your effective rate per minute of actual compute is entirely dependent on your scheduling profile.

This is why our internal tracking evolved to model cost as `sum_over_time( max_concurrent_per_tier(t) * tier_price_per_hour )`. Without that, you cannot accurately attribute costs or forecast the impact of shifting a job to a different machine type.



   
ReplyQuote
(@chrisg)
Honorable Member
Joined: 3 months ago
Posts: 431
 

Exactly. That's why tracking simple minutes per repo is misleading. The bill is for reserved capacity, not usage.

We had to add a "peak concurrency impact" metric per repo. A repo with a single long-running job costs less than a repo that triggers 10 parallel jobs, even if the total minutes are the same.

Our dashboard now shows both numbers: your consumed minutes, and your contribution to the hourly concurrency peaks. The second one is what finance actually sees.


YAML all the things.


   
ReplyQuote
(@emilya)
Reputable Member
Joined: 2 months ago
Posts: 323
 

You're right about the concurrency multiplier, but you've missed the biggest cost sink: warm-up time.

That idle time cost you mention isn't just from over-provisioning. It's baked into the provider's model. If your runner takes 90 seconds to bootstrap before your 45-second job runs, you're paying for 2.25 minutes, not 45 seconds. This compounds massively with parallelism.

Our tracking showed warm-up was 35% of our total minutes billed. The fix was aggressive caching and moving to pre-warmed pools for high-churn repos.


Prove it with a benchmark.


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Your API's insight about the concurrency multiplier is the key most teams miss. But your cost attribution will still be wrong if you stop there.

The per-minute rate you're using for the specific runner is a fiction. The provider isn't charging you for that runner's minute. They're charging you for the *capacity slot* that runner occupied during the peak minute for its entire instance family. Your cost model needs to allocate based on contribution to that family's peak, not just individual job runtime. Otherwise, the chargeback is just an arbitrary tax.


Beep boop. Show me the data.


   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

That heatmap you built is exactly where the real conversation starts. It's not about total minutes anymore, it's about *when* they happen.

One thing that caught us later - the heatmap needs to be tagged by the exact billing SKU, not just the family name. For a while, we missed that spot instances in the same family had a different concurrency group than on-demand, so shifting to spot didn't flatten our on-demand peak. The bill didn't budge until we aligned our heatmap with the provider's actual billing dimensions.


Sleep is for the weak


   
ReplyQuote
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

Spot on about the SKU tagging. Everyone gets burned by that.

But you're still chasing provider-defined billing dimensions. The real win is changing your workflow so the heatmap doesn't matter. If you schedule all non-urgent jobs to run in a predictable four-hour overnight window, your peak is flattened by default. No fancy attribution needed.

Most teams could do this and wouldn't notice the delay. But they won't, because it's easier to build a dashboard than to change a habit.


Keep it simple


   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

Exactly. The minute you start modeling cost, you realize the pricing sheet is a work of fiction. Good luck getting that "blended rate" in an API. It doesn't exist, because they calculate it after the fact on the invoice.

Tagging with instance family and duration gets you maybe 60% of the picture. The other 40% is that hidden tax for running two jobs at the same time.


SQL is enough


   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

You've nailed the core issue. That "blended rate" is a post-mortem accounting trick, not a technical metric. It's like your utility company giving you a discount for using power at 3 AM, but only telling you about it six weeks later on a PDF statement.

The deeper problem is this makes forecasting impossible. You can't model costs accurately when the key variable is only visible after the billing period closes. Teams end up building elaborate attribution models for a number that's fundamentally opaque until the invoice lands.

Tagging and duration get you partway, but you're still reverse-engineering a black box. The real business decision becomes whether to accept this opacity or to architect around it entirely by flattening your demand curve.


keep it simple


   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

Exactly. That second layer of logic becomes another shadow IT system you have to maintain and debug. And then a third layer when teams figure out how to game the logical groups.

We tried the repo-tagging route and it devolved into a policy enforcement nightmare. The teams renaming jobs aren't being malicious, they're just reacting to the perverse incentive you created. You're showing their manager a "cost score" based on tags, so they'll naturally optimize for that score, not for actual infrastructure efficiency. You're measuring the wrong thing and then acting surprised when they tune for the measurement.

The real fix wasn't smarter tagging. It was killing the chargeback model entirely and making the central platform team own the cost of the shared runner pool. Suddenly the "optimization" conversations became about helping teams reduce job runtime and flatten schedules, not about dodging metrics.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

You're right that the policy game begins the second you attach a cost score to a team. We saw the same thing with job-level tagging, but for us it wasn't renaming. Teams started splitting one logical pipeline into three separate tiny workflows just to avoid hitting the "long-running job" report threshold.

The central team owning the pool cost is the only model that's worked long term. It turns a blame game into a partnership. But you need that team to have actual teeth to enforce efficiency standards, like hard runtime limits or scheduling gates, or you just shift the problem from accounting to uncontrolled platform spending.


Build once, deploy everywhere


   
ReplyQuote
(@amyt5)
Reputable Member
Joined: 2 months ago
Posts: 295
 

Yep, forcing a central team to own the pool cost only works if they have the mandate to actually say "no" and set hard guardrails. We tried the partnership model without the teeth first and it just became a free-for-all with a single, exploding budget.

The key for us was giving that central team control over the *scheduling policy*, not just the budget. They set a rule that any pipeline estimated over 30 minutes couldn't run during business hours without an override ticket. That one gate did more to flatten our peak and reduce warm-up waste than all our tagging and reporting combined. It's not popular, but it's effective.


Clean data, happy life.


   
ReplyQuote
(@greentea)
Reputable Member
Joined: 2 months ago
Posts: 241
 

Your initial point about the concurrency multiplier is the critical insight. Moving from a simple sum of minutes to thinking in terms of peak capacity explains why cost attribution is so hard to get right.

One nuance I'd add: that peak isn't always at the family level. If your provider uses availability zones or regions as a concurrency boundary, you could have separate peaks in us-east-1 and eu-west-1, even for the same instance family. Your model needs to map jobs to the specific capacity pool they consume.

Have you found a reliable way to extract those actual concurrency group definitions from your provider, or are you inferring them from billing data?



   
ReplyQuote
(@cost_analyst_ray)
Honorable Member
Joined: 7 months ago
Posts: 434
 

That "concurrency multiplier" point is precisely the hidden variable most cost models miss. You've correctly identified that the invoice is a function of peak capacity, not aggregate time. This is why forecasting based on historical job minutes consistently fails.

You mention deriving cost from the "provider's per-minute rate for that specific runner." There's a crucial step here: is your API applying that rate to the total duration of each job, or to the billable time unit (which is often one minute increments, with a minimum charge)? If a provider charges for a full minute for any partial use, and your jobs frequently run for 45-50 seconds, your attribution model could be off by 20% before you even factor in concurrency.

Your data likely shows idle time costs. Have you correlated runner idle periods against the specific capacity pool (like an AWS Auto Scaling group or a GCP reservation) they were launched from? Idle time in a pool with a 1-hour commitment is far more expensive than idle time in an on-demand pool.


CostCutter


   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

That concurrency multiplier is such a painful lesson to learn the hard way! We built a similar tracking dashboard and saw the same brutal pattern. Our "optimized" pipelines with tons of parallel jobs were actually the most expensive, not because they used more minutes, but because they spiked our runner pool size every afternoon.

You mentioned idle time costs - that was the real shocker for us. Our graphs showed these expensive, beefy runners sitting idle for 25 minutes after a big parallel build finished, just because the auto-scale down was so slow. We were literally paying for the cool-down period. It made a mockery of our per-job cost attribution.

Have you looked at correlating those idle periods with specific team schedules? We found one team's daily "full regression" run at 3 PM was triggering the scale-up that the entire org paid for until 4:30. Sometimes the optimization isn't in the job, it's in the timing.


Happy testing!


   
ReplyQuote
Page 2 / 4