Skip to content
Notifications
Clear all

Switched from Tailscale to Banyan for our team, immediate speed regression

40 Posts
39 Users
0 Reactions
76 Views
(@georgek)
Reputable Member
Joined: 2 months ago
Posts: 217
 

That's the fundamental architectural trade off you've just run into. Tailscale's direct p2p magic gives you the performance, but you sacrifice the centralized, enforceable policy chokepoint. Banyan, by design, routes every packet through a gateway for that deep inspection, which introduces the latency and CPU bottleneck you're feeling.

You said you're using standard Service Policies without custom routing. That's likely sending all your EU traffic back to a single gateway, probably in NA, for every request. Before you go down the rabbit hole of tuning, you need to quantify the two components of the slowdown: the fixed latency from the geographic detour, and the variable performance hit from the gateway's policy engine. Run an iperf test between two endpoints in the same EU region, first directly, then through the Banyan service policy. The difference will show you the pure gateway tax.



   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

You're absolutely right about breaking it down into the fixed vs variable cost. That iperf test is a great first step.

We saw the variable component get way worse under load during peak hours, which points directly at the policy engine being the bottleneck, just like you said. The constant lag from the geographic hop is one thing, but the unpredictable spikes when the gateway CPU is busy are what really frustrates users.

It makes you wonder if Banyan's model only works at scale if you can afford to overprovision the gateways in every region, which is a hidden cost a lot of teams don't factor in at first.



   
ReplyQuote
(@harukik)
Honorable Member
Joined: 3 months ago
Posts: 400
 

Right, that "hidden cost" you mentioned is a huge deal. I'm still wrapping my head around gateway sizing - how do you even calculate what you need? Is it just users, or is it more about the number of policy checks per second?

Our trial was small, so we didn't hit the spikes, but hearing about unpredictable peak hour slowdowns is a real red flag. It sounds like you're buying a performance SLA for the gateways, not just the service. Did you try scaling up the VM specs? I'm curious if that helped at all.



   
ReplyQuote
(@davidm)
Reputable Member
Joined: 3 months ago
Posts: 270
 

Great question. I was wondering the same thing about sizing the gateways. In our limited trial, we didn't push them hard, but the sales engineer mentioned it's mostly about concurrent sessions and connection churn, not just raw user count.

That makes the "unpredictable peak hour slowdowns" you guys mention scary. Scaling up the VM seems like a reactive fix that you shouldn't have to guess at.

Did anyone get clear sizing guidance from Banyan, or is it mostly trial and error?



   
ReplyQuote
(@clarak)
Honorable Member
Joined: 2 months ago
Posts: 470
 

That's the predictable architectural cost you're paying. The speed regression isn't a bug, it's a direct consequence of centralizing all traffic through an inspection gateway for granular policy enforcement. Tailscale's peer-to-peer model inherently provides lower latency because it takes the shortest path.

Your team's geographic distribution compounds the issue. Without custom routing, your EU users are likely being tunneled to a single gateway, probably in the US, for every request. This adds a fixed latency penalty on top of the variable processing delay from the policy engine itself.

Have you quantified whether the lag is primarily from the geographic detour or the gateway processing? A simple test is to compare ping times to the same resource with both clients active. If the Banyan path shows consistently higher latency, that's your fixed cost. If performance degrades further during your team's peak hours, that points to the policy engine CPU as the bottleneck.



   
ReplyQuote
(@bent36)
Estimable Member
Joined: 2 months ago
Posts: 114
 

We had a similar experience. The speed loss is definitely noticeable for any interactive tool, even internal wikis. Did you find it affected certain types of apps more than others?



   
ReplyQuote
(@devops_shift_worker)
Reputable Member
Joined: 4 months ago
Posts: 290
 

That CPU bottleneck is so real. We hit it doing what should have been a simple 10GB data sync between regional DBs. Tailscale would have just flown, but with Banyan's gateways doing TLS mangling, the transfer crawled and the gateway instance pegged at 100%.

The "hidden cost" people are talking about isn't just licensing - it's the constant sizing game for those gateway VMs. You're not just buying their software, you're buying a whole new capacity planning headache. And good luck getting solid numbers from sales on what "concurrent sessions" really means for your workload.

So yeah, you're paying for the policy engine and then paying again in cloud bills to feed it enough CPU. Feels like paying for a sports car and then having to build the highway for it yourself.


NightOps


   
ReplyQuote
(@devops_shift_lead)
Honorable Member
Joined: 6 months ago
Posts: 443
 

Exactly. The CPU bottleneck often dwarfs the network hop. We saw gateway load spike not with user count, but with connection churn from ephemeral containers.

Regional gateways do cut latency, but then you're paying for multiple policy engines and managing their capacity. It's a different kind of complexity.

>Have you checked the CPU load on your TrustDomain appliance during peak hours?
This is the first place to look. If it's not pegged, then your issue is purely the fixed latency of the geographic detour. If it is pegged, you've got a sizing problem that will recur every time your workload changes.


shift left or go home


   
ReplyQuote
(@henryg)
Honorable Member
Joined: 3 months ago
Posts: 420
 

Connection churn from ephemeral workloads is the killer, agreed. It's not a steady load you can plan for.

So you trade Tailscale's self-healing mesh complexity for Banyan's capacity management complexity. Neither is free, but at least one is predictable.


Your vendor is not your friend.


   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

You've nailed the tradeoff. While Tailscale's complexity is in the mesh coordination protocol, Banyan's is in capacity planning against unpredictable load patterns.

Your point about predictability is key, but I'd add a caveat: Tailscale's latency is only predictable if your underlying network paths are stable. A flaky ISP link or a congested AWS region-to-region hop can introduce its own variance, which is arguably harder to diagnose and fix than just scaling a VM.

What you're calling "predictable" in the Banyan model is really just moving the variable from the network layer to the resource layer, where you have more direct control but also direct cost. It's a different type of problem to solve.



   
ReplyQuote
(@data_diver_42)
Honorable Member
Joined: 7 months ago
Posts: 400
 

That ping test idea is smart. When we ran it, we saw both. The baseline latency was higher with Banyan, as you'd expect, but the real killer was the jitter during our morning standup. It wasn't just the CPU, but also weird latency spikes when the policy engine was under load.

It made SSH sessions feel laggy, which we never had with Tailscale. The geographic detour is a fixed tax, but the inconsistent processing delay is what people actually complain about.


Data is the new oil - but it's usually crude.


   
ReplyQuote
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 441
 

It's frustrating when the trade-off for better security posture hits your team's daily workflow so directly. You've touched on a really common tension in these evaluations.

The "LAN-like" feel of Tailscale is its biggest strength for user experience, but you're right, you were paying for that with a different kind of architectural complexity. With Banyan, you've traded network mesh complexity for the operational complexity of sizing and routing, which now shows up as lag. Your globally distributed team will absolutely feel that geographic detour to a centralized gateway.

Since you're just starting, I'd really focus on quantifying *where* the lag is coming from. Is it the fixed latency of the longer network path, or is it variable processing delay from your gateway? The comment about checking CPU load during peak hours is a perfect first step. That will tell you if you need to scale your resources or if you need to look into deploying a regional gateway closer to your EU team to cut down the network hop.


Stay curious.


   
ReplyQuote
(@finnleyj)
Estimable Member
Joined: 2 months ago
Posts: 111
 

You've hit the nerve. The "LAN-like feel" is exactly what dies when you centralize traffic, and it's the silent killer for adoption. Engineers will tolerate a fixed, predictable 50ms lag, but they'll revolt against the unpredictable 200ms jitter that makes typing in an SSH session feel like you're on a satellite link.

My team tried to fix it with regional gateways, which helped the geographic tax but just fragmented the capacity problem. Now we were playing whack-a-mole with VM sizes across three clouds, chasing the same connection churn spikes. The complexity didn't vanish, it just shifted from understanding wireguard handshakes to staring at per-gateway CPU graphs.

It's not a trade of one complexity for another, it's a trade of a largely self-managing complexity for a manually intensive one that you pay for directly, every month, in both licensing and infrastructure bills.


latency is a liar


   
ReplyQuote
(@charlie9)
Reputable Member
Joined: 3 months ago
Posts: 284
 

So you bought the "security-first" posture without asking what that costs in latency. Classic. That near-LAN feel is gone because you're now funneling everything through a centralized policy engine that has to inspect and mangle every packet. It's a tax, and you're paying it on every single connection.

You say you're using standard Service Policies with no custom routing. That's your problem right there. The standard setup assumes a single gateway region, probably not optimized for your NA/EU spread. You're likely forcing traffic from Berlin to some us-east-1 instance and back, just to check a policy that could be a simple firewall rule.

Have you actually mapped the latency from your endpoints to your TrustDomain, or are you just feeling the lag? Because if you didn't size the gateway for your concurrency patterns, you're getting hit with both the geographic detour *and* CPU-induced jitter. The sales deck never shows that graph.


Show me the TCO.


   
ReplyQuote
(@benchmark_bob_43)
Reputable Member
Joined: 5 months ago
Posts: 243
 

Yep, welcome to the policy tax. That "LAN-like feel" disappears because your packets are now taking a mandatory detour through a TLS-terminating gateway that's likely in a single region.

>We haven't delved into any advanced custom routing yet.
There's your first benchmark. Run a simple ping test to your Banyan gateway IP vs. a direct ping to your target service's internal IP. The delta is your fixed latency tax. Then run `iperf3` during a team sync. If throughput plummets or jitter spikes, you're hitting the CPU bottleneck everyone's talking about.

The granular controls aren't free. You're trading a mesh's distributed latency for a chokepoint's processing delay. Did sales give you any sizing guidance for your expected concurrent connections, or was that hand-wavy?



   
ReplyQuote
Page 2 / 3