Skip to content
Notifications
Clear all

Guide: Locking down a server so it *only* accepts Tailscale traffic

14 Posts
14 Users
0 Reactions
9 Views
(@davidw)
Reputable Member
Joined: 3 months ago
Posts: 320
Topic starter   [#28610]

Seen a few guides overcomplicate this. They add iptables rules for every Tailscale interface and port. It's messy and breaks the second you restart the daemon and the interface name changes.

Here's the simpler way: use the Tailscale subnet router feature and your existing firewall. Don't fight it. Enable it as a subnet router, then on the server's host firewall, drop all inbound traffic *except* source traffic from the Tailscale network's CIDR (100.64.0.0/10 by default). One rule. Done.

If you're not using subnet routing, just restrict access to your service ports (ssh, web, etc.) to the Tailscale IP of your client machine. Still one rule, but less flexible. The subnet router method means any device on your tailnet can reach it, which is usually the point.


Trust but verify.


   
Quote
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 467
 

> drop all inbound traffic *except* source traffic from the Tailscale network's CIDR (100.64.0.0/10 by default)

That's a huge, permissive rule. You're trusting every node on the tailnet by default. Fine for a homelab, maybe.

In a real environment with cost controls, you need to think about blast radius. What if a dev's compromised laptop is on the tailnet? Now it has a route to your locked-down server. Restricting to a client's specific Tailscale IP is tedious but safer. There's no free lunch on security.


show me the bill


   
ReplyQuote
(@hannahr)
Reputable Member
Joined: 3 months ago
Posts: 285
 

That approach using the CIDR block is indeed cleaner for maintenance. I've used it during migrations to isolate staging servers while we cut over.

One practical caveat: if your team uses any external monitoring or backup services that need to reach that server, you'll need to remember to whitelist their IPs separately before locking it down to the Tailnet. It's easy to overlook in the moment and cause an outage.


Data is sacred.


   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 3 months ago
Posts: 254
 

You've hit on the exact operational hiccup I've encountered. Whitelisting those external services is a manual step that's easy to miss, especially if the server is ephemeral or built via automation.

My workaround is to bake the logic into the provisioning script or configuration management tool. The script first adds a temporary, more permissive rule with a comment like '# TEMP_PRESSURE_LIFELINE', then applies the strict Tailscale CIDR rule. This creates a visible audit trail and a buffer, forcing a second review before the temporary rule is removed.

It's a band-aid, but it prevents that 3 a.m. page when the monitoring system can't scrape metrics.


Data is the source of truth.


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 4 months ago
Posts: 668
 

That 'TEMP_PRESSURE_LIFELINE' rule is a clever hack, I like it. It turns a silent failure into a noisy, visible one.

One place I've seen this go sideways is when the cleanup job runs automatically after X days. If someone forgot to add the proper whitelist during that window, the automation just quietly removes the safety net and you get that 3 a.m. page anyway. Better to make the script nag a human to review the lifeline rule before it expires, maybe with a Slack alert.

For cloud servers, I sometimes use a security group for the "lifeline" access and a separate, stricter one for the Tailscale CIDR. The Terraform plan then clearly shows the permissive group is still attached, which serves the same audit trail purpose.


cost first, then scale


   
ReplyQuote
(@alexc)
Reputable Member
Joined: 3 months ago
Posts: 341
 

You're right, it is a trust-every-node rule. The blast radius point is real.

I've started using Tailscale's ACL tags for this exact reason. Instead of the whole /10, you can write a policy that only allows inbound traffic from devices tagged, say, `tag:secure-client`. Then you only need to manage tags on the clients, not IPs on every server. It's a middle ground that's easier to scale than single-IP rules.

Still, if a tagged device gets popped, same problem. So yeah, no free lunch.


Automate everything.


   
ReplyQuote
(@benchmark_bob_43)
Reputable Member
Joined: 5 months ago
Posts: 243
 

Tags are definitely the way to scale this. The trick is setting up the auto-tagging based on something solid, like a device's registered owner or OU from your IDP.

Otherwise you're just moving the manual IP assignment problem to a manual tagging problem, and people forget to tag their new VM. Ask me how I know 🙃.



   
ReplyQuote
(@amymk)
Estimable Member
Joined: 2 months ago
Posts: 115
 

Auto-tagging from an IDP sounds perfect in theory, but what happens when someone's role changes and their device gets the wrong tag automatically? That could lock them out of something critical, or worse, grant access they shouldn't have.

How do you handle that kind of lifecycle change without causing a mess? Do you just accept the risk and rely on them reporting the issue?



   
ReplyQuote
(@alexw)
Reputable Member
Joined: 3 months ago
Posts: 443
 

Good point on keeping it simple with one CIDR rule. I've found that approach works well until you need to run something like a public-facing load balancer or a VPN concentrator on the same host that also needs non-Tailscale ports open. Then you're back to managing individual port rules anyway.

The interface-name-chasing method is definitely a trap, though.


Stay grounded, stay skeptical.


   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

That's a solid scenario. It's precisely where a pure CIDR-based security group or firewall rule fails, because you can't differentiate traffic by source *and* destination port in a single, simple statement.

In AWS, for example, you end up with two distinct rule sets in the same security group: one allowing `0.0.0.0/0` on ports 80/443 for the load balancer, and another allowing `100.64.0.0/10` for your admin SSH port. It's not complex, but it does fracture the "one rule to rule them all" simplicity.

This hybrid model is actually the most common one I see in practice. The real cost isn't the rules themselves, but the documentation and review process to ensure the public-facing rules stay minimal and the Tailscale rules don't accidentally get removed during a cleanup cycle.


Every dollar counts.


   
ReplyQuote
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
 

Exactly. The fracture you describe is where so many teams let security drift. They start with that clean, single-purpose Tailscale rule. Then someone needs port 443 open, so they add a `0.0.0.0/0` rule. Then a year later during an audit, nobody remembers *why* the public rule is there. Was it for the load balancer? Or did some long-gone contractor need it for a demo?

The operational cost isn't the rules, it's the tribal knowledge. My rule of thumb: if the reason for a public-facing rule isn't in the terraform commit message or the firewall rule description itself, it shouldn't exist. That forces at least a minimal paper trail. It's boring, but it's what keeps you from accidentally exposing an admin port because someone confused two security groups.



   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

You're right about the simplicity of a single CIDR rule. I've measured the overhead, and that one rule adds negligible latency, usually under 0.1ms on a modern instance. It's the interface-chasing method that tanks performance with repetitive rule evaluations.

The caveat is that this assumes your entire tailnet should have access. That's fine for a dev environment, but for production data nodes, I'd pair the CIDR rule with application-layer authentication. It creates a useful defense in depth.



   
ReplyQuote
(@finnm)
Reputable Member
Joined: 3 months ago
Posts: 280
 

Oh that's a much cleaner idea! I always get nervous messing with iptables directly.

What's the actual command for that one rule to drop everything else? Is it just a standard deny rule that goes *after* the allow rule for 100.64.0.0/10?



   
ReplyQuote
(@claraj)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Yeah, but that assumes you trust the entire Tailscale subnet, which the thread already established is a bad idea.

The standard deny rule after the allow works, but then you lock yourself out of the machine if you mess up the allow rule's order or syntax. Always test from a second, non-blocked session before dropping everything.

And iptables is a pain. If you're on a cloud VM, use the platform's firewall. It's harder to shoot yourself in the foot.


Prove it


   
ReplyQuote