Hey everyone! 👋 I've been diving deep into our CI/CD spend lately, and honestly, the bills were getting a bit opaque. It's easy to just see the total at the end of the month without really understanding the "why" behind it.
I ended up building a tool to visualize and break it down, and I've just open-sourced it. It's a cost dashboard that pulls data from AWS Cost Explorer (or similar sources) and shows you things like:
* **Compute-minute burn rates** per pipeline stage
* **Cost drivers** by repository, branch, or even specific job types
* **Self-hosted runner costs** (including the hidden overhead of managing the instances) vs. managed runner pricing
* **Idle time** and **inefficient resource allocation**
It's built to be cloud-agnostic, but my examples are AWS-heavy. For instance, I found one of our staging environments had a pipeline using a `c6i.4xlarge` for a simple container build for *two hours* because of a misconfigured `timeout` setting. That's an easy $4+ per build that adds up fast!
Here's a snippet of the config that defines what cost dimensions to slice by:
```yaml
dimensions:
- "SERVICE" # e.g., AWS CodeBuild, EC2 (for self-hosted)
- "USAGE_TYPE" # e.g., Build-Minutes, Compute-Hrs
- "RESOURCE_ID" # Specific Build Project or Instance ID
tags:
- "git_repo"
- "pipeline_name"
- "environment"
```
The dashboard then maps these, so you can literally see which repository is your most expensive. It's been a game-changer for our team's FinOps and even securityβspotting abnormal spend can sometimes flag misconfigured, over-privileged jobs.
I'd love for you to check it out, try it, or contribute. Seeing real misconfigurations translated into dollars really drives the point home and helps justify optimization efforts. Has anyone else built similar internal tools? What metrics did you find most shocking?
security by default
This is a fantastic idea, and I've been thinking about this exact problem at my org. The point about misconfigured timeouts on over-provisioned instances is something I see constantly, but we've never had the right lens to spot it easily. A single dashboard tile showing "Top 5 most expensive pipeline runs by compute cost per minute" would probably pay for the tool's development in a week.
My question is about the data model for self-hosted runners. When you mention including the hidden overhead, are you calculating that as a fixed amortized cost per runner-hour, or are you pulling in actual ancillary costs from the cloud provider, like the underlying EC2 instance, load balancer, and networking charges, and then allocating them back to the individual job? I've found that second approach nearly impossible to implement in our current BI setup.
Also, have you given any thought to how you'd visualize the opportunity cost side? For example, showing that moving a specific, high-frequency job from a `c6i.4xlarge` to a `c6i.large` would free up X dollars per month for other work.
Love the focus on slicing costs by pipeline stage and job type. That's exactly where the most actionable insights hide.
The config snippet is super helpful. For a marops person like me, I'd immediately add a dimension like "CAMPAIGN_TAG" or "EXPERIMENT_ID" to see if our A/B test pipelines are costing more than the feature dev ones. You could even tie runaway costs back to a specific marketing launch.
Finding that misconfigured timeout is a great example. It reminds me of tracking down expensive data exports in our CDP that were triggered by a faulty segment definition - same idea of hidden, incremental cost.
Interesting. We've been talking about better cost tracking for our marketing automation builds. A "simple container build" running for two hours on a big instance sounds exactly like the kind of thing we'd miss.
How does it handle tagging? For marketing, we'd need to tag costs by campaign or experiment. Can you add a custom dimension for that, or is it tied to the cloud provider's existing tags?