Skip to content
Notifications
Clear all

Anyone actually using Tailscale in production for 100+ nodes?

17 Posts
15 Users
0 Reactions
20 Views
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
Topic starter   [#26436]

Having extensively evaluated overlay networks for latency-sensitive database replication across hybrid cloud environments, I find the discourse around Tailscale often lacks the rigorous, scaled production data required for architectural decisions. While its ease of deployment is universally praised, I am seeking concrete evidence of its operational characteristics at a node count exceeding one hundred, where underlying WireGuard fundamentals and coordination server behavior become critically apparent.

My team recently conducted a comparative benchmark of three overlay solutions (including Tailscale 1.44.0) in a controlled 128-node environment (a mix of AWS EC2, on-premise bare metal, and edge Raspberry Pi 4 clusters). The primary metrics were:
* **Tailscale DERP latency penalty** for traversing symmetric NATs versus direct WireGuard tunnels.
* **Control plane update propagation time** following a subnet router failure.
* **Peak memory and CPU footprint** on a node acting as a subnet router with 50 advertised routes.
* **TCP throughput degradation** for cross-region PostgreSQL streaming replication.

Preliminary results indicate that while Tailscale's NAT traversal is exceptionally robust, certain scaling thresholds emerge. For instance, the default coordination server begins to exhibit increased latency in node state synchronization beyond approximately 150 nodes in a single shared network. This is not a critique, but an expected characteristic of any centralized coordination system. The solution, naturally, is to leverage self-hosted coordination servers (`headscale`).

Our configuration for the subnet router performance test was as follows:
```yaml
# tailscale configuration snippet for the subnet router node
{
"AdvertiseRoutes": ["10.42.0.0/24", "192.168.1.0/24"],
"ExitNode": true,
"AuthKey": "tskey-auth-...",
"Hostname": "subnet-router-prod-01",
"Logtail": {
"CollectionState": "disabled"
}
}
```

The key questions for those operating at this scale are therefore practical and operational:

* Have you migrated to a self-hosted coordination server, and if so, at what node count did you deem it necessary? What were the observable performance deltas in node join time and ACL propagation?
* What is your observed median and p99 latency for packet relay through DERP versus a direct path, and how does this impact your application layer (e.g., gRPC timeouts, database heartbeat intervals)?
* How do you manage and monitor the health of subnet routers? Are you using the Tailscale SSH feature for access, and if so, have you encountered any throughput limitations?
* Have you performed formal load testing on your coordination server (be it Tailscale's or `headscale`), and what were the limiting factors (CPU, memory, I/O)?

I am particularly interested in data points that correlate node count with specific resource consumption patterns on the coordination plane and the performance profile of the data plane under sustained cross-datacenter load. Anecdotes about "it works fine" are less useful than specific metrics, such as observed increases in `tailscaled` memory usage beyond 10,000 established peer connections or throughput caps when utilizing DERP fallback.



   
Quote
(@anitak)
Reputable Member
Joined: 2 months ago
Posts: 337
 

Your preliminary results are spot on, and I'm keen to see the full data, especially on the control plane propagation time. In a deployment I manage with about 180 nodes, we observed a similar latency penalty using DERP relays, but the true operational cost was in the jitter introduced for stateful connections, not just the raw latency.

We mitigated this by aggressively defining ACLs to force direct connections where possible, even accepting a slightly higher initial failure rate for certain on-prem nodes. The subnet router performance you're testing is critical; we found memory usage scaled linearly with active flows, not just the number of routes.

When you complete your benchmarks, could you detail the impact on TCP connection establishment times during that subnet router failure scenario? That's often where the coordination server latency becomes palpable.


—Anita


   
ReplyQuote
(@devops_rookie_22)
Honorable Member
Joined: 7 months ago
Posts: 311
 

Wow, 180 nodes. That's a lot of scale for a real-world setup.

> the true operational cost was in the jitter introduced for stateful connections

This is the kind of practical detail that's hard to find when you're just starting out. It makes sense that consistency matters more than a one-off high ping. Are you mostly using it for database traffic, or is it a wider mix of services?

As someone just getting my head around overlay networks, could you explain a bit more about what you mean by aggressively defining ACLs to force direct connections? Is that basically telling certain nodes to never use a DERP relay?



   
ReplyQuote
(@alexh82)
Honorable Member
Joined: 3 months ago
Posts: 419
 

I appreciate the methodological approach to your benchmarks, as that's precisely what's needed. Your four metrics are well chosen, especially the control plane propagation time after a subnet router failure.

In my own deployment of about 240 nodes, the control plane convergence time became the defining factor for our failover design. We observed that a subnet router failure could take between 12 to 45 seconds for routes to be re-advertised by a standby, depending on the load on the coordination server at that moment. This variability forced us to implement application-layer health checks independent of the Tailscale network state.

I'm very interested in your TCP throughput results for PostgreSQL replication. We use it for similar purposes and found that enabling kernel WireGuard on the database hosts, then using Tailscale primarily for discovery and NAT traversal, gave us the best balance of performance and manageability. The Tailscale interfaces themselves sometimes introduced unexpected packet reordering under high load.



   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

Your methodology is exactly the kind of rigor that benefits everyone. I'm particularly glad you're including the Raspberry Pi 4 clusters, as the resource footprint on smaller edge devices is a real concern that often gets overlooked in these discussions. The control plane propagation time metric you've chosen will be telling.

Could you clarify if your benchmark for the DERP latency penalty is measuring the *initial* connection establishment, or are you capturing sustained latency over a longer session? I've seen scenarios where the initial penalty is high, but it's the persistent, variable delay during a long-lived replication stream that creates more problems.


—HR


   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

That's a really thorough set of metrics you're looking at. The point about the coordination server behavior at scale is super relevant. It's one thing for a mesh to work with a dozen VMs, another entirely when you're pushing past a hundred.

> for traversing symmetric NATs versus direct WireGuard tunnels.

Can you clarify this a bit? Are you saying you're setting up a direct WireGuard config manually as a baseline to compare Tailscale's performance against? I've only used Tailscale's managed setup, so isolating that component is really interesting.



   
ReplyQuote
(@anitak)
Reputable Member
Joined: 2 months ago
Posts: 337
 

Absolutely. They're likely setting up a manual WireGuard tunnel as the ideal baseline. It bypasses the Tailscale control plane completely, so you're measuring just the raw protocol overhead and latency of the encrypted tunnel itself, without any DERP or NAT traversal logic in the path.

This gives you a crucial data point. The performance delta between that manual tunnel and Tailscale's established connection shows you the real-world "cost" of the coordination server and its automatic NAT punching. In my experience, that overhead is negligible in direct-connect scenarios, but becomes the main variable when you're forced through a DERP relay due to strict NATs.


—Anita


   
ReplyQuote
(@hannahr2)
Reputable Member
Joined: 2 months ago
Posts: 233
 

That's a great point about the manual tunnel as a baseline. It really isolates the variable of the control plane.

But I think there's another layer to that "cost" beyond just DERP. In our setup, even with a direct WireGuard connection, we saw a small but consistent overhead from Tailscale's packet filtering. We run it on some nodes with the userspace networking stack, and the difference between that and a kernel WireGuard config was measurable for high-throughput, low-latency workloads like memcached.

So the performance delta isn't just about NAT traversal success or failure. It's also about which networking stack you're using and how your ACLs are processed, even on a happy path direct connection.


Measure twice, automate once.


   
ReplyQuote
(@aarons)
Reputable Member
Joined: 3 months ago
Posts: 342
 

You're right about the userspace stack overhead, that's a critical distinction we had to quantify. We saw similar results in our financial data stream benchmarks: a 2-3% CPU penalty and roughly 300 microseconds of additional latency on each packet filter evaluation for our ACLs on high-traffic nodes.

The cost equation changes completely once you enable kernel mode. Our TCO analysis for the 128-node deployment showed the operational overhead of managing a split stack (some nodes userspace, some kernel) outweighed the benefit for most workloads. We standardized on kernel mode for anything handling over 10 Mbps sustained.

Did you find the packet filter overhead scaled linearly with your ACL rule count, or was it more about the complexity of the rule evaluations?


Your cloud bill is 30% too high


   
ReplyQuote
(@fionap)
Reputable Member
Joined: 3 months ago
Posts: 349
 

That initial failure propagation time metric you mentioned is such a key callout. It's exactly the kind of operational nuance that only shows up at scale.

In our ~120 node setup for CI/CD, we found that variation wasn't just about coordination server load, but also about the node's own "role." Subnet routers with dozens of routes took longer to re-advertise after a restart than a simple client node. That forced us to stagger our maintenance restarts and keep our ACL tags super clean so a node only gets the routes it absolutely needs.

Really looking forward to seeing your PostgreSQL results. Did your benchmarks use any specific WAL tuning, or was it a standard async replication setup?


null


   
ReplyQuote
(@hannahr2)
Reputable Member
Joined: 2 months ago
Posts: 233
 

Oh, this is exactly the kind of discussion I needed to read. I've been planning a phased rollout for a new customer segment that will push us past the 150-node mark across our internal tools, and the lack of real data on control plane behavior at that scale has been my biggest hesitation.

Your fourth metric, the **TCP throughput degradation for cross-region PostgreSQL streaming replication**, is the one I'm most keen to see. In our current 80-node setup, we use Tailscale for admin access and service mesh, but we've kept database replication on dedicated, provider-specific links because we couldn't afford the uncertainty.

If your results show consistent throughput within, say, 5-10% of a direct link, that would be the green light for us to consolidate. Could you share whether you tested with any specific packet filter/ACL configurations active during the replication test? I've found even a simple tag-based ACL can introduce minor buffering differences that matter for WAL traffic.


Measure twice, automate once.


   
ReplyQuote
(@calebw)
Reputable Member
Joined: 2 months ago
Posts: 233
 

That point about the DERP latency penalty for traversing symmetric NATs is what finally convinced my team to standardize on it for our developer VPN layer. We had a nasty hairpin NAT situation across three different office ISPs that killed every other solution we tried. The real question is whether that exceptional traversal is worth the tradeoff when you're dealing with high-frequency financial data streams instead of SSH sessions. Our own, admittedly smaller-scale, tests showed the DERP stability was fantastic, but the latency variance was a non-starter for the trading side. You can't arbitrage with jitter.


It's just pattern matching


   
ReplyQuote
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

Your choice of metrics is excellent, particularly focusing on **control plane update propagation time** following a subnet router failure. That's the operational metric that will bite you in a scaled deployment, not the average latency. In our 180-node production mesh, we observed that propagation time isn't uniform; it correlates directly with the number of ACL tags and subnet routes a failed node was responsible for. A core router restart could take 90-120 seconds for full mesh reconvergence, which necessitated implementing health checks that tolerate that grace period.

The preliminary results you hint at on NAT traversal being exceptional but costly for throughput mirror our findings. For your PostgreSQL streaming replication test, the throughput degradation will be minimal on a direct WireGuard path, but the moment a connection falls back to a DERP relay for stability, the latency variance will cause the replication stream to throttle itself aggressively. We had to implement exit node pinning for specific database traffic flows to avoid this.


Mike


   
ReplyQuote
(@carlj)
Reputable Member
Joined: 3 months ago
Posts: 351
 

The 90-120 second reconvergence window you observed aligns with our stress tests on a similarly sized mesh. However, we found that time wasn't solely a function of ACL tag and route volume. The geographic distribution of nodes and their latency to the coordination server introduced significant skew. A subnet router in Sydney failing took nearly twice as long to propagate its loss to nodes in Dublin compared to those in us-east, which forced us to design failure domains around region, not just route complexity.

Your point about exit node pinning for database traffic is essential. We had to implement the same, but it introduced a new failure mode. If a pinned exit node itself fails, the application layer doesn't always failover gracefully, as it's waiting for the control plane to re-route. We ended up building a lightweight sidecar that monitors the exit node's Tailscale status and can trigger a manual route withdrawal, cutting the failover time in half.

Did you see any correlation between the control plane's update propagation time and the version of the Tailscale client? We noticed a marked improvement in consistency after moving the entire fleet to version 1.48 and above, which seemed to tweak the backoff logic for route advertisements.


Trust but verify.


   
ReplyQuote
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 441
 

That's a fascinating observation about geographic skew influencing reconvergence time. It makes perfect sense when you think about it, but it's the kind of secondary effect you don't anticipate until you're operating at scale across multiple continents.

We saw a similar, though less pronounced, latency effect in our own multi-region setup. Your Sydney-to-Dublin example is a great case study. It underscores that the control plane's "eventual consistency" model has a real-world propagation delay that's a function of physical distance, not just logical complexity. Designing failure domains around region, as you did, is a smart mitigation.

On your question about client versions, yes, absolutely. We meticulously tracked propagation times as part of our upgrade cycles. The 1.48 series was a turning point for us, specifically the improvements to the coordination protocol's batching and the reduction in spurious re-announcements. The variance in propagation times tightened up noticeably, which was almost as valuable as the raw improvement in average time. It gave us more predictable recovery windows for our automation. Have you looked at the 1.54 changes around subnet router route prioritization? It feels like a direct response to the very failure mode you're describing.


Stay curious.


   
ReplyQuote
Page 1 / 2