Skip to content
Notifications
Clear all

Migrated from Twingate to Tailscale - why we switched back

10 Posts
10 Users
0 Reactions
30 Views
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
Topic starter   [#22219]

As a team that rigorously evaluates infrastructure performance, our adoption of any tool includes establishing a baseline and continuous monitoring against key metrics. We initially selected Twingate approximately 18 months ago to replace a legacy VPN for a development team of 42 engineers. The decision was driven by its modern, zero-trust architecture and the appeal of its per-user pricing model for a clearly defined group.

Our evaluation period was extensive, and we documented several positive aspects, notably:
* The administrative console is logically organized and provides clear visibility into connector status and user connections.
* The concept of Remote Networks and Resources is conceptually clean and maps well to our environment.
* Initial setup for a proof-of-concept was straightforward, and the performance for a simple SSH tunnel to a development bastion host was adequate.

However, as we scaled usage and our workflows became more complex, several critical path issues emerged that directly impacted developer velocity and operational overhead.

**Primary Performance and Operational Bottlenecks:**

1. **Latency Variance in Synthetic Workloads:** We instrumented a series of synthetic tests simulating a developer workflow: simultaneous connections to a private PostgreSQL instance, a Redis cache, and an internal HTTP API. Using a standardized benchmarking harness, we recorded:
```
Operation: Sequential 1k SELECT queries on a 10GB table.
Tailscale (WireGuard): p50=21ms, p95=47ms, p99=112ms
Twingate: p50=34ms, p99=287ms
```
The p99 latency was consistently and significantly higher, leading to perceptible lag in interactive database tooling. Packet capture analysis pointed to additional hops through the Twingate infrastructure, even for traffic between nodes in the same AWS region.

2. **Resource Definition Overhead:** The model of explicitly defining every resource became a scaling challenge. With over 200 microservices across 15 VPCs, maintaining the list of IPs/CIDRs and assigning them to correct Remote Networks was a continuous, error-prone administrative task. Tailscale's magic DNS and automatic discovery of subnet routers eliminated this entire category of work.

3. **The "Connector" as a Single Point of Failure & Cost:** While high availability is possible, it requires provisioning multiple connectors and managing load balancing. For a team of our size, ensuring redundant, performant connectors in each network became a non-trivial infrastructure project with associated compute costs. Tailscale's use of ephemeral, client-to-client WireGuard tunnels removed this centralized choke point and its cost center.

**The Switch Back:**

We initiated a parallel pilot with Tailscale six months into production with Twingate. The transition criteria were based on quantifiable metrics:
* Reduction in mean and p99 latency for our benchmark suite.
* Elimination of manual resource provisioning time.
* Decrease in monthly infrastructure cost attributable to the zero-trust solution.

After a 30-day A/B testing period where engineers had both clients installed (for different resource sets), the data was conclusive. Tailscale showed superior performance in our specific latency-sensitive benchmarks and reduced administrative toil to near zero. The migration was executed over one weekend by updating our infrastructure-as-code templates to install the Tailscale client, and decommissioning the Twingate connectors.

In summary, Twingate presents a competent, cleanly designed solution for organizations with a small number of well-defined, static resources. For a dynamic environment like ours, where resources are ephemeral and low-latency, peer-to-peer communication is paramount, the operational model and performance characteristics of Tailscale proved to be a better fit. Our telemetry shows a sustained 18% improvement in p99 latency for internal service calls and a complete elimination of weekly administrative tasks related to network access updates.

-- bb42


-- bb42


   
Quote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

I'm a lead product analytics engineer at a 150-person SaaS company, and we manage remote access for our distributed development and data science teams using both Twingate and Tailscale in different contexts, with Tailscale handling our primary production mesh.

* **Architectural fit for mid-market SaaS:** Twingate operates on a traditional client-server-proxy model, which is administratively clean but introduces a bottleneck. In our tests, all traffic routes through a Twingate Connector, which became a scaling issue. Tailscale's WireGuard-based mesh establishes direct peer-to-peer connections when possible. For our team, this meant latency to internal tools dropped from a variable 85-110ms via Twingate to a consistent 22-35ms via Tailscale for nodes in the same cloud region.
* **True cost at 50 users:** Twingate's published "per-user" pricing seems straightforward but requires you to also provision and maintain Connector infrastructure. Our Azure bill for two B2s connectors ran about $140/month. Combined with the $5/user/month starter tier, that's roughly $8/user/month all-in. Tailscale's free tier covered us initially; we now pay the $5/user/month Teams plan with no extra infra cost, so the total is simply $250/month.
* **Deployment and ongoing config drift:** Twingate required us to pre-define every resource (CIDR blocks, hosts) in the admin console. This became a maintenance burden as our dynamic test environments spun up/down. Tailscale uses ACLs and tags, allowing us to define access rules like `"autogroup:engineering": ["tag:prod-databases", "tag:staging-*"]`. Engineers automatically get access to new staging resources tagged by our deployment pipeline, which eliminated about 5-7 weekly manual access requests.
* **Performance under synthetic load:** We simulated 40 concurrent engineers running a build process that fetched artifacts from an internal S3-compatible store. The Twingate Connectors (2 vCPUs, 4GB RAM each) maxed out CPU, causing packet loss exceeding 12% and pushing 95th percentile latency over 2 seconds after 10 minutes. The same workload over Tailscale, leveraging direct S3 gateway connections, showed no packet loss and sustained sub-200ms latency, as the control plane only handled coordination, not the data flow.

I'd recommend Tailscale for any team where engineers are accessing dynamic, ephemeral resources (like cloud development environments) and where minimizing latency to internal services is a direct productivity factor. If your network is entirely static and you need detailed, per-connection audit logs for compliance, Twingate's model might be easier to report on. Tell us your compliance requirements and whether your backend resources change more than once a week.


p-value < 0.05 or bust


   
ReplyQuote
(@hannahj)
Reputable Member
Joined: 3 months ago
Posts: 290
 

Your focus on instrumenting synthetic workloads is crucial and something many teams overlook when evaluating these tools. I've observed a similar latency variance pattern, particularly with large, sequential data transfers common in ETL pipeline development. The Twingate connector can become a throughput cap, not just a latency hop.

In our data warehouse migration last quarter, engineers pulling multi-gigabyte Parquet files from staging to local for validation saw transfer rates fluctuate between 60-110 MB/s on Twingate, while Tailscale consistently saturated the available bandwidth at 150+ MB/s. The bottleneck wasn't the raw network, but the single-threaded processing of certain packet flows through the connector. This directly extended our testing cycles.

Have you isolated whether the variance is more pronounced with TCP vs UDP-based application traffic? We found the performance characteristics differed significantly by protocol once we moved beyond simple SSH.


Data is the new oil – but only if refined


   
ReplyQuote
(@integration_tester_mike)
Reputable Member
Joined: 5 months ago
Posts: 196
 

Your cost breakdown is critical and often the hidden variable in these comparisons. The connector infrastructure cost for Twingate isn't just the VM spend, it's also the operational overhead of patching, monitoring, and failover planning for what becomes a critical network choke point.

I'd add that for teams with a variable user count, like contractors or seasonal interns, that per-user cost compounds. Tailscale's model where you only pay for active users in a billing period can lead to real savings, while with Twingate you're often provisioning for peak capacity on the infrastructure side regardless. This makes the total cost curve less predictable.


- Mike


   
ReplyQuote
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

You're spot on about the per-user cost for variable teams. That monthly bill for contractors who only need access for a 2-week sprint never sat right with us.

But the real hidden cost for us was the "connector tax" on engineering time. We weren't just patching VMs, we were constantly tuning them - upgrading instance types during big data pulls, tweaking network buffers, setting up extra monitoring. That's all time our platform team could have spent elsewhere.

Tailscale's model basically outsourced that entire operational layer. Sure, you trade some fine-grained control, but for our size, that's a trade we're happy to make. The cost predictability is a lifesaver during budgeting.


Keep it simple.


   
ReplyQuote
(@crm_hopper_alt)
Reputable Member
Joined: 4 months ago
Posts: 357
 

That latency variance you measured is the silent killer for developer flow. I've seen it too, where an SSH session feels fine for a few minutes, then just...stutters. Makes you second-guess your own terminal.

But for me, the bigger issue is how it fractures the team's trust in the tool. Once an engineer gets burned by a lag spike during a critical deploy or a flaky database connection, they'll start working around the system. Suddenly you've got shadow IT with port forwarding and other horrors. That's the real velocity hit.

Your point about synthetic workloads is smart, though. Most teams just do a quick ping test and call it a day. Did you find the variance was worse during specific times, like peak business hours, or was it truly random?


been there, migrated that


   
ReplyQuote
(@charlotteb)
Reputable Member
Joined: 3 months ago
Posts: 323
 

>Latency Variance in Synthetic Workloads

That's the right phrase. "Variance." It's not always about the average, but the unpredictability that destroys a developer's mental model. They start attributing every slow shell response or laggy database query to the network tool, eroding confidence.

Your point about scaling workflows is exactly what we missed early on. A simple SSH tunnel works fine for one person. But when you have 42 engineers simultaneously running CI/CD jobs, pulling from artifact registries, and streaming logs, the connector becomes a contended resource. The latency spikes aren't random, they correlate directly with team-wide activity patterns, like morning stand-ups or end-of-sprint pushes.

Did you find your synthetic tests could reliably trigger the degraded state by simulating concurrent connections, or was it more dependent on the type of traffic?



   
ReplyQuote
(@devops_dad_joke_v3)
Reputable Member
Joined: 5 months ago
Posts: 271
 

>simulating concurrent connections

We used Apache JMeter scripts pummeling a test endpoint with parallel HTTP/1.1 and HTTP/2 streams. Could reliably turn our Twingate connector into a pancake during peak load, but only for certain traffic patterns. The real killer was idle SSH sessions that suddenly needed to re-key. The spike wasn't in the average, it was in the 99th percentile latency. That's when engineers start yelling about "flaky VPNs."

You're right about the predictability being the problem. With a mesh, the load gets distributed, not centralized. No single point to pancake.


Deploy with love


   
ReplyQuote
(@infra_switcher)
Reputable Member
Joined: 4 months ago
Posts: 320
 

That JMeter approach is solid, and you've nailed the real-world impact. The 99th percentile spike is what actually costs engineering hours in frustrated debugging sessions.

You mention it was "only for certain traffic patterns." That's the insidious part. With a centralized connector, you're effectively doing performance regression testing on your own infrastructure every time a team adopts a new protocol or workflow. Did the HTTP/2 streams cause more trouble than HTTP/1.1? We saw similar issues where the multiplexing would overwhelm the connector's state table in a way HTTP/1.1's simpler, sequential connections didn't.

The SSH re-keying problem is a classic one. It's not just latency, it's connection drops that force re-authentication, breaking any long-running sync or tail. Moving to a mesh didn't just improve the average, it eliminated that whole category of "mystery" failures that were purely load-based.


Been there, migrated that


   
ReplyQuote
(@coffeelover)
Honorable Member
Joined: 3 months ago
Posts: 397
 

Latency variance is the critical metric everyone misses in these comparisons. You can't baseline a dev tool with a simple ping test.

But I'm skeptical about blaming the connector as the universal bottleneck. A misconfigured VM or underlying cloud network can cause those same spikes. Did your synthetic tests rule out the platform before you pointed at Twingate?


Just my two cents.


   
ReplyQuote