Alright, let's settle this. Everyone's got opinions on CI/CD migrations, usually about "developer experience" or "ecosystem synergy" or whatever buzzword is floating around this week. I don't trust feelings. I trust milliseconds and dollars.
So when the directive came down from on high to move our 20-odd service pipelines from TeamCity to GitHub Actions, my team groaned. Another "modernization" project. I decided if we were going to waste engineering cycles on this, we were at least going to measure *everything*. Built a simple dashboard to track the only metrics that matter: build duration, queue time, and compute cost per pipeline, before and after the cutover.
The setup was straightforward, if tedious. A scrappy Postgres table to store timestamped pipeline run data, scraped from both platforms' APIs. A couple of views to normalize the data. Grafana on top. The real pain wasn't the dashboard—it was the translation.
Here's a taste of what "pipeline translation" actually meant. Taking a declarative, mature TeamCity configuration and turning it into a YAML soup of `actions/checkout@v4` and third-party actions that might break tomorrow.
**TeamCity (The sane way):**
```kotlin
steps {
script {
name = "Build and Test"
scriptContent = """
docker build -t myapp:${build.number} .
docker run --rm myapp:${build.number} pytest /app/tests
""".trimIndent()
}
dockerCommand {
name = "Push"
commandType = push {
namesAndTags = "myregistry.com/myapp:${build.number}"
}
}
}
```
**GitHub Actions (The YAML carnival):**
```yaml
jobs:
build:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Build and Test
run: |
docker build -t myapp:${{ github.run_number }} .
docker run --rm myapp:${{ github.run_number }} pytest /app/tests
- name: Push to Registry
uses: docker/build-push-action@v5
with:
push: true
tags: myregistry.com/myapp:${{ github.run_number }}
```
Looks simpler? Wait until you need to manage shared secrets across 20 repos, or implement a coherent rollback strategy, or debug why your job is stuck in queue for 10 minutes because the `ubuntu-latest` runner is oversubscribed.
The results after a month of running both systems in parallel? Inconclusive, which is its own kind of result. For 14 of the 20 services, the average build time was within +/- 5%. Three got slower (mostly due to colder caches and longer queue times on the GitHub runners). Three got faster (lightweight Node.js services that benefited from the simpler setup). The real cost delta was in the ops overhead—managing our own TeamCity agents versus paying for GitHub's minutes.
The moral of the story isn't that one platform is better. It's that a migration like this is rarely about performance. It's about consolidating the vendor list, or following the herd, or making a manager's slide deck look good. If you're going to do it, at least measure the actual impact. Otherwise, you're just trading one set of quirks for another, newer, shinier set.
Dashboard snippet for the curious: simple table with service name, 7-day avg build time (old), 7-day avg build time (new), delta, and delta percentage. The most dramatic change was a 12% increase for our monolithic API service. Took us a week to tune the caching strategy to get it back to baseline.
-- old salt
Exactly, this is the core problem everyone glosses over. The translation cost from a mature system to a YAML-based workflow is enormous and never reflected in those "5% faster build" benchmarks. You're not just moving configs, you're rebuilding institutional knowledge as code.
Your example is about to show a clean, versioned Kotlin DSL, and the GHA equivalent will be 50 lines of brittle steps. Did you track the metric that matters most: pipeline config maintenance time? My team found that for complex pipelines, GHA increased the time to implement a change by a factor of three, because you're debugging opaque action behaviors instead of logic you own.
The dashboard is good, but I'd add a layer for tracking infrastructure drift. With TeamCity, you control the agents. With GitHub, you're at the mercy of their runner images and network. I've seen builds break because a default Python version changed on `ubuntu-latest` overnight.
—davidr
Yep, the hidden maintenance cost is real. But you're still missing the biggest line item: the bill.
> With TeamCity, you control the agents. With GitHub, you're at the mercy...
You're at the mercy of their pricing, too. Self-hosted runners are free, but now that's your hardware and ops. The managed runners? The minute you need a bigger machine or longer runtime, the cost balloons. Those "free" minutes disappear fast. Did your team factor that "infrastructure drift" into actual dollars, or just stability?
You're right to point out the bill, but it's more predictable than you think if you treat it like any other vendor contract. The real trap isn't the cost ballooning, it's the auto-scaling. You set up a matrix build or accidentally trigger a workflow on a branch push, and you've burned through a month of minutes in an afternoon because the platform happily lets you.
Your procurement team should be negotiating commit tiers with upfront pricing, not paying as-you-go. Otherwise, you're right, you're just trading hardware ops for financial surprise ops.
Show me the data
> I trust milliseconds and dollars.
Good. That dashboard is the only way to prove the ROI, or lack of it. Everyone else just argues in the abstract.
What was the data pattern? For us, queue time dropped to near zero with GHA's scale, but average build duration increased by 15% due to cold starts and less predictable agent hardware. The cost per build looked better on paper until we added the engineering months spent rewriting and debugging those YAML workflows.
Did you see a clear winner, or just a different set of trade-offs?
Your point about cold starts and agent hardware variability matches our observations almost exactly. We recorded a 12% median duration increase, with a much wider variance - the 90th percentile build time increased by nearly 30%, which caused significant planning disruption.
The different set of trade-offs is the critical finding. The "winner" depends entirely on your organizational bottlenecks. If you were agent-constrained in TeamCity, the queue time elimination is transformative. If your builds were already efficient and stable, you're trading a known, fixed cost for variable performance and a new financial model.
We attempted to quantify the rewrite cost as you suggested, translating it into a per-build amortization over three years. It made the cost-per-build comparison markedly worse for the first 18 months, which is a timeline most business cases ignore.
Migrate slow, validate fast.
Yeah, that variance in build times is the killer. It makes planning a nightmare, even if the median looks okay. We saw something similar, but with spot instance preemptions on a self-hosted k8s runner setup. The 99th percentile times were wild, way beyond the 90th.
Amortizing the rewrite cost is a great idea, but painful to see. Most teams just call it a "one-time migration" and move on, ignoring that it's a real cost that should be factored in. Did you find any config patterns that helped reduce that variance, or is it just an inherent tax on managed runners?
Self-host or die trying.
Exactly the right approach. That translation pain is where the real story is. A clean, versioned Kotlin DSL versus a sprawling YAML file full of black-box actions.
One thing I'd add: the *ongoing* translation. It's not a one-time cost. Every time you need to update a pipeline because a third-party action changes its interface or gets deprecated, you're back in YAML soup, debugging something you don't own. We started a simple registry of "trusted" actions, but it's still a new source of friction the old system didn't have.
Did you standardize on any specific patterns to try and lock down that YAML variability?
Automate the boring stuff.
Oh, the 99th percentile preemption spikes sound brutal. We were on managed runners, so our variance came from cold starts and the underlying Azure hardware lottery. We did find a couple of config patterns that helped flatten the worst outliers.
First, we started using the `actions/cache` with a much more aggressive key strategy. If a cache hit could skip even a minute of setup, it offset some cold start penalty. Second, we explicitly avoided the `ubuntu-latest` label in critical paths and pinned to a specific runner image like `ubuntu-22.04`. That removed the variable of an in-place OS upgrade mid-build.
But you're right, it's mostly an inherent tax. The real "pattern" was a cultural shift: we stopped expecting deterministic build times and started designing for it. We broke one massive pipeline into three separate, faster jobs that could run concurrently, so a single slow runner didn't block the whole show.
Did your spot instance setup let you define any anti-affinity rules to reduce the chance of getting preempted mid-matrix?
Integration Ian
Ah, the "scrappy" Postgres table and Grafana dashboard. Classic. Did you also track the "dollars" part, or just the milliseconds? The moment you push to production, you're on the clock with GitHub's billing. That's where your nice graphs go to die.
Read the contract
The translation from a true declarative DSL to YAML is the unacknowledged cognitive tax. You're trading a typed, versioned configuration for a stringly-typed workflow that's essentially a series of remote procedure calls to black-box binaries.
You mentioned the `actions/cache` pattern in a later post, which helps, but it's a workaround for a fundamental problem: you're rebuilding your environment on every run. In TeamCity, the agent state is a managed entity. With GHA, you start from a vanilla image and hope the network to fetch dependencies isn't having a bad day. That adds unpredictable latency that your median might hide but your p99 will expose.
Did you track cache hit rates as a separate metric? It becomes a critical performance lever, turning your build into a distributed systems problem.
--perf
You mentioned normalizing the data between platforms. How did you handle cost attribution for the GHA runs? The pricing isn't as direct as per-minute compute, especially if you used a mix of public repos or private minutes. Did you have to estimate it?
We didn't normalize cost.
We gave up. It was impossible to get an apples-to-apples number because GHA's billing model is opaque and blended. You get a monthly invoice for "minutes," but what's the underlying compute unit cost? Nobody knows.
We just used the invoice total and divided by builds. It's inaccurate, but so is any other method. The real cost is the team constantly trying to decipher it.
Simplicity is the ultimate sophistication
The cognitive load of translating from a true DSL to YAML is immense and underreported. It's not just about syntax; it's a paradigm shift from a controlled, typed environment to string concatenation masquerading as configuration. Your Kotlin example is spot on - it enforces structure.
The most insidious part is the loss of local reasoning. In TeamCity, a step is a known quantity. In the YAML soup, each `uses:` line is a call to an opaque, versioned binary. You now have a dependency graph of external actions, each with its own failure modes and update schedules, buried in your pipeline definition. Did you find a way to version-lock these third-party actions effectively, or does it just become a periodic audit burden?
Your point about pinning the runner image is crucial. We found the `ubuntu-latest` label was a huge source of silent variance; it could shift from Ubuntu 20.04 to 22.04 between queued jobs on the same day, introducing subtle toolchain differences. Pinning to a minor version like `ubuntu-22.04` gave us a stable baseline for dependency hashing, which made our aggressive `actions/cache` keys actually reliable.
On the anti-affinity question for spot instances: no, not in a useful way. The cloud provider's spot market algorithms are a black box. We tried spreading a matrix across availability zones, but the preemption decisions seemed random at our scale. The only effective strategy was treating every job as potentially interruptible, which meant designing for idempotence and checkpointing artifact state much earlier in the pipeline. It turned our build scripts into distributed systems code.