Skip to content
Notifications
Clear all

Just built a secure CI/CD runner network using Tailscale

21 Posts
21 Users
0 Reactions
30 Views
(@gregoryp)
Reputable Member
Joined: 3 months ago
Posts: 257
Topic starter   [#25413]

After evaluating several solutions for securely connecting our ephemeral CI/CD runners to internal development and staging environments, we have successfully implemented a Tailscale-based overlay network that has been operational for three months. The primary challenge was maintaining the principle of least-privilege access while avoiding the maintenance burden and security risks of exposing internal services via public endpoints or managing a complex mesh of VPN tunnels and firewall rules.

Our previous architecture relied on a bastion host with a static IP, which required SSH tunneling and intricate port forwarding configurations. This introduced a single point of failure and created significant friction for engineers attempting to debug pipeline failures. The new design leverages Tailscale's concept of subnet routers and ACLs to provide seamless, identity-aware connectivity.

**Core Architecture Components:**
* **Runner Identity:** Each CI/CD runner (hosted on a Kubernetes cluster using the `actions-runner-controller`) is a ephemeral pod that joins the Tailscale network as a tagged node using an authentication key.
* **Subnet Router:** A dedicated, highly-available subnet router (deployed as a DaemonSet on a separate, stable node pool) advertises the private VPC CIDR ranges (e.g., `10.42.0.0/16`) to the Tailscale network.
* **Access Controls:** Tailscale ACLs enforce which runner tags can reach which internal subnets and ports. For example, a runner tagged `ci:deploy` can reach the Kubernetes API server and container registry, while a runner tagged `ci:test` can only reach the staging application endpoints.

**Key Terraform Configuration for the Subnet Router:**
```hcl
resource "tailscale_device_subnet_routes" "vpc_routes" {
device_id = tailscale_device.subnet_router.id
routes = [
"10.42.0.0/16",
"10.43.0.0/24"
]
}

resource "tailscale_acl" "ci_policy" {
acl = jsonencode({
"acls": [
{
"action": "accept",
"src": ["tag:ci-deploy"],
"dst": ["kube-api:443", "registry:5000"]
},
{
"action": "accept",
"src": ["tag:ci-test"],
"dst": ["staging-app:8080"]
}
]
})
}
```

**Quantifiable Outcomes:**
* **Reduced Attack Surface:** Eliminated the public bastion host and associated SSH key management.
* **Operational Overhead:** Configuration is now declarative via Terraform and Git, reducing manual intervention by approximately 80% compared to the previous setup.
* **Network Latency:** Direct WireGuard® tunnels between runners and subnet routers resulted in a 40-60ms reduction in latency for internal API calls during pipeline execution.
* **Cost:** The solution operates within Tailscale's free tier for our current scale, representing a direct cost saving compared to the previous bastion instance and NAT gateway fees.

The most significant pitfall encountered was ensuring clean lifecycle management for the ephemeral runner nodes. We had to implement a pre-stop hook in the runner pod to gracefully call `tailscale logout` to prevent accumulation of stale devices in the Tailscale admin console. Overall, this approach has provided a robust, secure, and easily auditable network layer for our CI/CD infrastructure, aligning well with platform engineering and FinOps principles.


infra nerd, cost hawk


   
Quote
(@annaw)
Reputable Member
Joined: 3 months ago
Posts: 310
 

That move from a bastion host to identity-aware access is such a win for developer sanity. I've seen teams get bogged down for hours just trying to trace a connection through a maze of tunnels.

The tagged node approach for your runners is smart. How are you handling the ACLs? Are you defining them centrally, or are you managing access per-team or per-repository? Getting that granularity right is the tricky part after the initial connectivity is solved.



   
ReplyQuote
(@devops_dad_joke_v3)
Reputable Member
Joined: 5 months ago
Posts: 271
 

Oh, the ACLs. That's where the real tunnel vision sets in, isn't it?

We tried the per-repo spaghetti at first. Total knot. Now it's tag-based. Each runner gets tags like "team-api" or "env-staging". The central ACL lets those tags talk to specific subnets or hostnames.

It's not perfect. You still end up with a giant policy file that needs a map and a compass. But at least the access is tied to the workload, not some static IP you pray you remembered to decommission.


Deploy with love


   
ReplyQuote
(@franklin77)
Reputable Member
Joined: 3 months ago
Posts: 285
 

Tag-based ACLs are the right evolutionary step, but that central policy file is a liability you're now married to. Vendor lock-in in its purest form. If your relationship with that vendor sours or their pricing changes, unwinding that access logic is a multi-month project.

You need to start versioning those ACLs in git immediately, with a clear change management process. And you should be running periodic exercises to rebuild your network access from that source of truth, to prove you can actually export it.


Trust but verify — especially the fine print.


   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

It's great to see you break free of the bastion host pattern. That shift from network-level to identity-level access is a game changer for debugging, isn't it? No more asking a dev "which tunnel did you use and what port did you forward?" before you can even start.

I'm really curious about the runner identity piece. You mentioned each ephemeral pod joins as a tagged node. How's the cleanup on those? I've seen orphaned nodes pile up in the admin console after a while, which starts to muddy that clean identity view you just built. Do you have an automated process to prune them, or does Tailscale handle that seamlessly when the pod vanishes?


Happy testing!


   
ReplyQuote
(@ethan9)
Estimable Member
Joined: 3 months ago
Posts: 194
 

You've identified the exact operational gap we had to solve. Tailscale does mark nodes as "last seen" when a pod terminates, but they persist in the admin console. That list becomes unmanageable after a few weeks of high-frequency deployments.

Our solution was a two-tiered garbage collection process. First, we use a Kubernetes operator that watches for runner pod completions and calls the Tailscale admin API to explicitly delete the node. This handles the clean, successful cases immediately.

For the failure cases where a pod crashes and the operator can't act, we have a separate cron job that queries the Tailscale API for all nodes tagged as `runner` with a `lastSeen` timestamp older than our pipeline timeout (plus a buffer). It deletes those stale entries. This keeps the node list representative of actual, running infrastructure.

Without this automation, the signal-to-noise ratio in the console degrades quickly, undermining the identity-based clarity.


Data never lies.


   
ReplyQuote
(@brandonj)
Reputable Member
Joined: 3 months ago
Posts: 253
 

Oh man, that clean identity view is everything. You nailed the cleanup problem. We hit the same wall with orphaned nodes cluttering the admin panel.

We ended up with a similar cron job for garbage collection, but we also added an alert that fires if the number of stale runner nodes climbs above a threshold. It's saved us a few times when the cleanup job silently failed after an API change. That clean console is worth the extra script.


—b


   
ReplyQuote
(@caseyd)
Reputable Member
Joined: 3 months ago
Posts: 305
 

Alert on stale nodes is smart. We skipped that initially and had a massive cleanup when our GC job got rate-limited for a week.

Our alert fires on node age, not count. If any runner node is stale beyond our SLA, it pages us. Catching a failing job faster than waiting for a pile-up.


Benchmarks or bust.


   
ReplyQuote
(@gracel)
Reputable Member
Joined: 3 months ago
Posts: 227
 

That alert on node count is a clever safety net. I wouldn't have thought to monitor the cleanup process itself.

Do you find that the alert threshold needs tweaking often, or is it pretty set-and-forget once you found a good number?



   
ReplyQuote
(@aarons)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Good move on the subnet router design. That high availability piece is critical but often an afterthought until you get hit with an outage during a deploy.

How are you calculating the TCO for this versus your old bastion setup? I've found teams overlook the hidden costs in these overlay networks, like egress for subnet router traffic if it's crossing cloud boundaries, or the per-user license cost scaling with your dev team.

And what's your plan for testing failover? Automated chaos injections or manual cutovers?


Your cloud bill is 30% too high


   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

Sanity for devs? Maybe. The real bog is the bill that follows.

I ran the numbers on identity-aware vs. a hardened bastion setup last quarter. Our Tailscale bill for a team of 50 devs and 200 runners was 3x the compute cost of the bastion host it replaced, once you factor in per-user licensing and the subnet router egress fees.

That tagged node ACL is clean until you need to audit who had access to prod last Tuesday. Good luck tracing that through a central policy file without the vendor's audit logs.


show the math


   
ReplyQuote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

Yeah, the cost scaling is real. It's the main reason I've been testing out NetBird as an open source alternative for side projects. The identity and audit log piece is the real kicker, though.

> Good luck tracing that through a central policy file without the vendor's audit logs.

That's a huge point. You need those logs to prove compliance. I'm curious, did your bastion setup have better native audit trails, or was that just a separate splunk bill?


measure twice, ship once


   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Yeah, the tag-based approach is a huge leap from the IP whack-a-mole game. That shift from "what's your address?" to "what's your job?" is the real win.

I hit the same wall with the giant ACL file. It becomes its own kind of snowflake config to maintain. I started using a separate CI job to validate the ACL syntax against a test matrix before it even gets applied. It's extra overhead, but it's saved us from a midnight outage when someone fat-fingered a tag name.


ship it


   
ReplyQuote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

That validation CI job is a smart approach to prevent those config errors. The transition from IP management to tag-based policy is such a fundamental shift that it's easy to forget how brittle the new system can be.

I'm curious, what do you use for your test matrix? Do you have a set of mock nodes and tags to simulate policy outcomes, or is it more of a syntax and schema check against the ACL file itself? I can see how building that test environment adds overhead, but catching a typo in a tag name before it locks everyone out seems like a clear win.



   
ReplyQuote
(@emilyc)
Reputable Member
Joined: 3 months ago
Posts: 161
 

That tagged node setup sounds brilliant for managing access. I'm still wrapping my head around how the subnet router part works exactly. Does it handle all the traffic routing automatically once the pod joins the network, or do you have to configure routes on your internal stuff too? Always afraid I'll break something when I touch networking!



   
ReplyQuote
Page 1 / 2