Hi everyone,
First, I want to thank this community for all the insights I've gathered here over the past few months. Your posts were a huge help when I was planning this migration. As the title says, I recently finished moving our company's fleet of just over 200 development and testing devices—a mix of cloud VMs, on-prem workstations, and container hosts—from a traditional VPN setup to Tailscale. It's been mostly smooth, but I definitely learned a few things the hard way.
The biggest win has been the simplicity of access control. Using tags and ACLs to define which groups of devices can talk to each other has been a game-changer for our CI/CD pipelines. Instead of managing a complex mesh of firewall rules, we can now just tag a new staging server and our GitHub Actions runners can reach it instantly. That said, my mistake was not planning our tag naming schema carefully enough at the start. We ended up with a messy mix of `env:prod`, `prod`, and `production` tags before we standardized. A bit of upfront design would have saved us a migration later.
Another lesson was around subnet routers. We have a legacy physical network for some hardware, and I initially set up every possible device as a subnet router "just in case." This created unnecessary complexity and some routing loops during a brief outage. I've since learned to be minimalistic: only declare subnet routers where they are absolutely needed, and document the traffic flow for each.
On the containerization side, running the Tailscale container in our Kubernetes clusters was straightforward, but we hit a minor snag with the `--accept-routes` flag not being set correctly in our initial Helm chart, which isolated some pods. A quick config tweak fixed it.
Overall, the move has drastically reduced our support tickets for "I can't connect to X." For anyone managing a similar number of devices, my advice would be to start with a solid ACL and tagging strategy on paper, go slow with subnet routers, and leverage the `tailscale` CLI for bulk operations. It's a powerful tool.
Happy to answer any questions based on our experience!
still learning
That point about tag naming schema really resonates. We had the same mess with client environments, bouncing between `client:acme`, `acme`, and `env-acme`. We ended up enforcing a simple `client_name:environment` prefix rule (like `acme:prod` and `acme:staging`) across the team. It made writing those ACLs so much cleaner.
How did you handle tagging the container hosts versus the workloads running on them? That's where our logic got a bit tangled.
Comparing tools one review at a time.
Interesting timing, I was just looking at Tailscale for a similar sized project. Your point about tag planning is well taken, we've had that problem with user groups in other tools.
You stopped mid-sentence at the end there - what happened with the subnet routers? Did you scale back to just a few, and how did you decide which ones?
The subnet router question is a critical one. Our initial plan was to deploy them on every major on-prem network segment for "redundancy". That created a mesh of advertised routes that was a nightmare to debug and introduced multiple single points of failure into the routing layer itself.
We scaled back to just two strategic locations per physical site, chosen based on network topology, not device count. The primary is a small, dedicated VM on the core network with high uptime. The secondary is on a different hypervisor cluster. The decision criteria were simple:
- The host must have a stable, low-latency path to the site's default gateway.
- It must be a managed node, not a developer workstation.
- We avoid placing them on devices that handle other high-throughput workloads to prevent resource contention.
This reduced the advertised route table complexity and let us treat the subnet routers as infrastructure, not as part of the application fleet. The key lesson was that subnet router redundancy is about placement quality, not quantity.
--perf