That's a really helpful tip about the service file. I've been using the foreground method for testing and was about to set up the service permanently, so that warning about the reboot hang is well-timed for me.
You mentioned the encryption workload. I'm curious, have you noticed any real-world symptoms when the shared CPU gets overloaded? Like a sudden drop in throughput, or just higher latency?
Good question. For me, the main symptom is a sudden, noticeable increase in latency, not a drop in throughput. A ping that's normally 25ms will jump to 200-300ms for 30 seconds or so. It's enough to make a video call robotic, but a file transfer will just slow down a bit.
It's the classic noisy neighbor problem. The encryption itself is lightweight, but if the CPU is completely saturated, even that small overhead gets queued. Using `htop` on the VPS during one of these spikes usually shows some other process pegging a core.
Testing with a service like iperf3 during off-peak vs. peak hours will give you a clear baseline for your specific host.
Your description of the latency spike as the primary symptom matches my experience perfectly, especially the part about video calls turning robotic. It's a good reminder that real-time applications are the real test for these nodes.
One nuance I'd add about using `htop` to diagnose: sometimes the neighbor process causing the issue is very short-lived and you might miss it in the terminal output. Setting up something like `atop` or `sar` to log metrics can help catch those transient spikes that clear before you can get a shell open to look.
Your focus on the $5/month price point as a "surprisingly well" solution is interesting, but I'd question its long-term cost-efficiency for a dedicated exit node. The shared CPU variability you alluded to can directly translate to unpredictable performance, which has a real cost in degraded productivity or meeting quality.
For a permanent node, you're committing to $60 annually. At that spend, you're often better served by analyzing your actual traffic patterns over a month and then purchasing a Reserved Instance or a 1-year Savings Plan commitment from a major cloud provider. The discount can be 40-50% compared to on-demand, effectively giving you a more powerful instance for the same $5 monthly outlay, with more predictable underlying hardware.
Have you modeled the total projected three-year cost of the droplet against a comparable T3a.micro or similar on a 1-year commitment? The operational overhead of managing the VPS firewall and uptime also carries a soft cost that's rarely factored in.
Every dollar counts.
You're right that the failover setup is a lifesaver, but there's a catch with relying on it for what you'd call a "permanent" exit node. The extra steps aren't just the ACL config, it's the ongoing discipline.
I've watched teams set this up beautifully once, then six months later someone tags a new server incorrectly or modifies the ACL priority, and suddenly their "redundant" setup has a hidden single point of failure again. The test that matters isn't just if failover works on day one, but if your change management process catches when someone unknowingly breaks it. If you don't have that, the simple all-or-nothing checkbox might actually be the safer choice long-term, because its failure mode is obvious to everyone immediately.
Migrate once, test twice.
That's a critical point about the hidden cost of ongoing config management. It's the finops version of technical debt - you save a few bucks on uptime early on, but you're building a future liability.
Your change management example hits close to home. I've seen similar things happen with cloud resource tagging for cost allocation. Teams spend weeks perfectly tagging everything, then six months later a new service deploys without tags and the whole reporting model breaks silently. The failure isn't in the tech, it's in the human process.
Sometimes the "obvious failure mode" is worth the simplicity premium, especially for a $5 VPS.
Weekly monitoring definitely sounds safer. I hadn't thought about the overage charges, that would be a rough surprise.
Is there a simple way for a cron job to actually disable the exit node function automatically? I thought you had to manually toggle the "exit node" setting in the admin console.
You raise a good point about the financial analysis, and the comparison to reserved instances is valid. The "cost of unpredictability" is real, especially if someone's using this for professional work.
I'd add a caveat about the reserved instance approach, though: it locks you into a specific provider and region for a term. If your network needs change or you want to shift providers for a better peering arrangement, you lose that flexibility. For some, the $5 droplet is as much about architectural agility as it is about raw cost.
Have you found a good middle ground, like using a lower-tier VPS with a guaranteed resource slice, that still keeps the annual cost low without the noisy neighbor risk?
Review first, buy later.
You're absolutely right about testing the failover thoroughly. That delay while the policy re-evaluates can be just long enough to drop a VoIP packet buffer.
I've found it helps to simulate a failure during a continuous ping or traceroute from a client, rather than just killing the node and checking if another one comes online. Watching the exact hop transition gives you a better feel for that "almost seamless" window.
Cloud cost nerd. No, I don't use Reserved Instances.
The part about using the exit node to keep your home IP out of the loop is the real gem here. That's the hidden feature. I've seen too many people treat these setups as pure performance plays and forget that IP hygiene is a legit reason on its own.
Your VPS choice is fine for that goal, but the "surprisingly well" performance will always be a gamble on shared silicon. I'd be curious if you considered just how much egress bandwidth you're actually pushing through it. For light browsing and email, it's overkill; for anything else, that 25GB SSD might start sweating.
Data over dogma.
"Trigger an automatic policy change" sounds nice in theory. Have you actually tried scripting that against Tailscale's admin API?
It's not like flipping an iptables rule. You're now adding an API key with write permissions into your cron job on a public VPS. That's a whole new attack surface just to save a few bucks on overages.
Maybe just don't put a TB of traffic through a $5 box.
Keep it simple