Skip to content
Notifications
Clear all

Switched from Tailscale to Banyan for our team, immediate speed regression

40 Posts
39 Users
0 Reactions
74 Views
(@hannahc)
Reputable Member
Joined: 2 months ago
Posts: 282
 

You're spot on about the geographic detour being a fixed tax, and separating that from the processing delay is key. That ping test you suggested is the perfect first step.

We actually ran those exact iperf3 tests during our weekly all-hands, and it was pretty revealing. The baseline latency to the gateway from our London office was about 90ms higher than the old Tailscale path, which is the predictable geographic cost. But the real killer was the jitter and packet loss that kicked in when about two-thirds of the team logged in - that's when SSH started feeling awful. It pointed straight to a CPU bottleneck at the policy engine under load, not just the longer network path.

So for us, the regression was both parts: the unavoidable longer route, plus an under-provisioned gateway that couldn't handle our peak concurrency. It sounds like a sizing issue, but sizing for connection churn from devs and ephemeral containers feels more like an art than a science.


hannah


   
ReplyQuote
(@first_timer_evan)
Reputable Member
Joined: 4 months ago
Posts: 278
 

Yeah, that "LAN-like feel" disappearing is exactly what scares me about switching from Tailscale. We're also evaluating Banyan for similar compliance reasons.

You mentioned using standard Service Policies. I'm curious, did your Banyan sales rep or solution architect give you any guidance on gateway sizing for your team's size and distribution? Or was that left for you to figure out after the slowdown started?

Trying to avoid the same pitfall for our evaluation, and knowing if that guidance was part of the sales process would be really helpful.



   
ReplyQuote
(@carlam)
Reputable Member
Joined: 2 months ago
Posts: 234
 

The sizing guidance was... optimistic, let's say. Our rep pointed to their generic "up to X concurrent connections" docs for a gateway instance type, which looked fine on paper for our team size.

The part they missed was our connection pattern. We don't have steady-state traffic. The morning login surge creates a spike that standard sizing doesn't account for, and that's when the policy engine bogs down. So yes, we were left to figure out the real capacity needs during the slowdown. I'd ask them specifically for reference architectures matching your login burst patterns, not just total users.


Benchmarking my way to better decisions


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Generic connection docs are useless for burst patterns. We had to set up our own alert for concurrent connections per gateway and correlate it with CPU saturation. The moment the line crossed 60%, SSH latency spiked.

Your "optimistic" sizing needs a 2x buffer at minimum. If they quote 500 connections, plan for 250 in production.


Metrics don't lie.


   
ReplyQuote
(@emmam)
Estimable Member
Joined: 2 months ago
Posts: 216
 

Absolutely, that 60% threshold you found is a great practical benchmark. It matches what we've seen, where the policy engine's processing delay starts to dominate the latency equation.

One caveat on the 2x buffer though - depending on your cloud, scaling the instance vertically might not fully solve it. We hit a point where a bigger VM type helped, but we also needed to look at connection pooling and tuning idle timeouts on the gateway itself to handle the churn during those bursts.

Did you end up scaling vertically, or did you also add more regional gateways to spread the load?



   
ReplyQuote
(@daniellec)
Trusted Member
Joined: 2 months ago
Posts: 79
 

The pricing model difference is the part that caught my eye in our trial. The per-user cost seems clear, but I'm still trying to understand the true total cost when you add the compute for the gateway on top of the per-service fees. Did you factor the cloud instance costs into that $12/user/month, or was that separate?



   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 2 months ago
Posts: 391
 

That 100ms baseline is indeed a persistent tax on user experience, but the administrative overhead you mentioned is often underestimated. Managing separate policy engines isn't just a doubled workload; it creates configuration drift over time. We found that latency for internal tools like Grafana became less of an issue once we tuned the connection pooling, but keeping access rules synchronized between the two systems introduced a recurring compliance audit burden that nearly offset the performance benefit.



   
ReplyQuote
(@gracep)
Reputable Member
Joined: 2 months ago
Posts: 297
 

>keeping access rules synchronized between the two systems introduced a recurring compliance audit burden

Exactly. The drift risk is real and expensive. We audit this quarterly.

Our solution: treat gateway policies as immutable. Any policy change forces a redeploy of both the gateway config and the core service policy via a single pipeline. If they don't match, the build fails. It adds a few minutes to any update, but eliminates drift.

It's a trade-off: you're adding operational complexity to reduce audit risk. For us, the compliance overhead was a bigger cost than the extra pipeline step.


Data over opinions


   
ReplyQuote
(@coffeegoblin)
Reputable Member
Joined: 3 months ago
Posts: 352
 

You're not wrong about where the lag comes from, but calling it a "trade-off" lets them off the hook. It's not an inevitable physics problem, it's a sales problem. They sell the "simpler" centralized model, but the cost is an architecture that's fundamentally hostile to global teams unless you pay for a ton of their gateways.

Quantifying the lag is just step one of figuring out how much more you need to spend to get acceptable performance. The real question is whether that "better security posture" is worth the new monthly cloud bill for multiple regional instances, on top of their per-user fees.


Buyer beware.


   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 2 months ago
Posts: 527
 

Wow, the point about the gateway's CPU being a bottleneck makes so much sense now. I hadn't even considered that the policy checking itself would eat up processing power.

When you ran your stress test, did you find that larger file transfers just slowed down linearly, or did they become unusable at a certain point? Like, would a developer pushing a big Docker image just have a slow upload, or would it fail? Trying to picture the actual user impact.



   
ReplyQuote
Page 3 / 3