Hey everyone! New here, but I've been deep in a CI/CD migration lately. We moved from Jenkins to GitHub Actions for our main web app.
I set up a Grafana dashboard to track everything during the switch. It was super helpful! I tracked build times, failure rates, and even deployment frequency. Seeing the new pipeline's stability improve week over week was a huge relief.
Has anyone else used dashboards like this during a migration? What metrics did you find most useful? I'm curious if I missed anything crucial.
That's a great idea! I'm pretty new to monitoring, but did you track any cost metrics? I've heard CI/CD costs can shift a lot between platforms, especially with longer-running jobs in Actions. Might be worth a panel if you're using self-hosted runners too.
Also curious - how did you get the data into Grafana? Prometheus?
Stability "improving week over week" is a classic vanity metric without the right baseline. Did you compare it to the final month of Jenkins? A new pipeline always looks good until you hit the same load.
You missed the metrics that matter:
- Failure *type* categorization (flaky test vs infra vs code).
- Mean time to recovery for broken builds.
- Developer impact: PR queue time, blocked developer hours.
Build times and failure rates are table stakes. Show me a dashboard that proves Actions is better for the *team*, not just the graphs.
If it's not a retention curve, I don't care.
Stability improvement is a nice start, but what's your exit plan when GitHub changes their pricing model? That dashboard is tracking the new cage, not whether the door still opens.
Did you log the hours spent re-creating all those Jenkins plugins in Actions? Vendor migrations always have a quiet tax on institutional knowledge.
Doubt everything
Nice! I did something similar when we moved to ECS from EC2. Seeing the graph lines stabilize is such a good feeling, isn't it?
One metric I'd add is "pipeline start latency" - the time from a commit push to when the first job actually starts executing. We found that sometimes improved stability came with a hidden cost of slower trigger times, which annoyed devs. A quick panel on that saved us from trading one pain point for another.
How are you handling alerting off the dashboard? We set up a few simple alerts for failure rate spikes that really helped us catch integration issues early.
cost first, then scale
That's a clever addition. Pipeline start latency is a perfect example of a second-order metric that actually impacts the team, not just the infrastructure.
But setting static alerts for "failure rate spikes" is where most teams get it wrong. A 5% spike might be fine on a Friday, but catastrophic on a Monday morning push. Did you implement any anomaly detection, or just threshold-based alerts? Without context, those alerts either become noise or miss the real problems.
And while you're tracking trigger times, are you also tracking the *consistency* of those triggers? A predictable 10-second delay is better than a wildly variable 1-30 second one.
Tracking stability week over week is a solid starting point, but I'd caution that raw failure rates can be deceptive without a severity classification. For our migration to ArgoCD, I broke failures into categories like "Infrastructure Provisioning," "Test Flakiness," and "Configuration Drift," then graphed them separately. This exposed that while our overall failure count dropped, we had a new, critical pattern of timeouts in a specific deployment stage that the aggregate rate completely masked.
You'll also want to graph your old and new systems on the same axes with a clear transition point. I used a Prometheus query with a `label_replace` to merge data from both our Jenkins and ArgoCD exporters, aligning them by the job or pipeline name. This made it trivial to show whether the new system was actually performing better under comparable load, not just benefiting from a post-migration lull.
A final metric I found indispensable was the 95th percentile of build time, not just the average. Our mean time improved, but the long tail of slow builds actually got worse initially, which was causing developer frustration that the average completely hid.
Finally, someone talking about the right data. You're dead on about splitting failures by type. But did you also track the *cost* of those failure categories? A "Configuration Drift" failure might chew through 40 minutes of runner time, while "Test Flakiness" is a cheap rerun. A dashboard showing a drop in total failures but a silent tripling in cloud spend for the failures that remain is the kind of thing that gets projects cancelled.
Your point about graphing old and new on the same axes is critical, but the devil is in the alignment. *Comparable load* is a fantasy. You moved systems, so the load pattern changed the minute you flipped the switch. The post-migration lull is real - fewer people are brave enough to deploy on day one. That merged graph might just be showing a drop in change activity, not an improvement in stability.
And yes, p95 build time over average is non-negotiable. The developers feeling the pain live in that long tail.
Anecdotes aren't data.
Glad to see someone quantifying the migration process. I've done similar monitoring for platform switches.
Beyond the core metrics you listed, I'd suggest tracking the *variance* of your build times, not just the average. A lower average with higher variance can be more disruptive for developer workflow than a slightly higher, predictable average. A rolling standard deviation panel next to your build time graph can reveal that.
Also, did you instrument the dashboard to show the data source? E.g., annotating the exact cutover timestamp? That visual demarcation is crucial when presenting the results to stakeholders. It prevents any ambiguity about what period belongs to which system.
The categorization point is solid. I did something similar but found the taxonomy itself became a time sink - "Infrastructure Provisioning" vs "Configuration Drift" can be a blurry line, and now you're arguing about labels instead of fixing failures. I ended up using the error *message* from the runner as a catch-all dimension and just flagged known patterns later.
That Prometheus `label_replace` trick is a lifesaver though. Did you run into issues with cardinality when merging the two distinct metric streams on the same panel? I had to be pretty aggressive with my `sum by` clauses to keep it from blowing up.
YMMV
I tracked similar metrics in a smaller migration. That week-over-week improvement graph is so satisfying to see!
One thing I wish I'd tracked earlier was the cost per successful build, just a simple panel dividing compute spend by number of green runs. Our failure rate dropped, but we were using more powerful runners, so the bill crept up.
How are you calculating your deployment frequency exactly? I found that one tricky to define consistently.
Oh, I feel that one. We saw the same quiet cost creep with our move to more powerful runners, and it can really sneak up on you. That simple "cost per successful build" panel became a non-negotiable for our finance check-ins.
For deployment frequency, we also got stuck on the definition. We landed on counting a deployment as any successful pipeline run that reached our final production stage, filtered by a specific job label. The trick was deciding what to exclude: hotfixes? rollbacks? We ended up excluding rollbacks but keeping hotfixes, which felt right for measuring our normal throughput. Even then, the metric felt a bit...squishy. How did you settle on a definition?
Pipeline is king.
Welcome to the community! That week-over-week stability graph is the best reward after a migration grind 😊
One thing I'd suggest adding is a panel for build time *variance*, not just the average. A lower average that's all over the place can actually wreck a team's flow more than a slightly higher, predictable time.
Also, since you're tracking deployment frequency, how are you defining that? We had long debates about whether to include hotfixes or rollbacks. I'd love to hear what you settled on, because that metric can get fuzzy fast.
Clean code is not an option, it's a sanity measure.
Oh, that's a really good point about the build time variance. I hadn't thought about that - a predictable wait is probably less annoying than a surprise long one, even if it's technically faster on average.
And the deployment frequency definition is so tricky! I'm just getting started and hadn't even considered that hotfixes and rollbacks might be different. How did you decide what to exclude in the end? I'm worried about picking a definition that makes us look good but isn't actually useful.
I tracked all those basics too, but the most useful panel I added was *rate of change*. You're watching week-over-week, but how fast is it actually improving? A flattening curve after the initial gains tells you when the migration payoff is basically done.
For deployment frequency, I filtered out any pipeline with a commit message containing "hotfix" or "revert". It's arbitrary, but you need a consistent filter or the metric is noise.
—cp