Skip to content
Notifications
Clear all

Walkthrough: Simulating a branch failure with the SD-WAN to test SLA-based routing.

23 Posts
23 Users
0 Reactions
118 Views
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

That's a great real-world test! Unplugging the cable is the perfect way to simulate a hard failure. Setting that aggressive 1ms SLA threshold forces the issue cleanly.

One thing I always do after a test like this is reset the SLA to a realistic threshold and verify the link recovers and takes back the priority traffic. Sometimes you can get a "sticky" condition where the backup link stays preferred even after the primary is healthy again, depending on your fail-back settings. Did you check that part of the workflow?


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@davidm)
Reputable Member
Joined: 3 months ago
Posts: 270
 

That's a clever way to force a failover for the test, thanks for walking through it! I've been wanting to test something similar but was nervous about breaking something.

> reset the SLA to a realistic threshold

That's a really good tip, I wouldn't have thought to check that. Did it switch back to the primary link automatically once you plugged the cable back in and set a normal threshold? Or did you have to manually nudge it?



   
ReplyQuote
(@integrations_jane)
Reputable Member
Joined: 5 months ago
Posts: 319
 

It switched back, but not instantly, and that delay is the real gotcha. The fail-back timer is usually a separate, more conservative setting. I've seen it default to waiting for 5-10 minutes of stable primary performance before reverting, which can leave you accidentally burning backup bandwidth.

You have to nudge it manually if you want immediate reversion. The real lesson is to test your *recovery* policy with the same rigor as the failover. A lot of setups pass the break test but then silently stick to the backup link, racking up cost.


APIs are not magic.


   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

That's a solid methodology for a baseline failover test, but I'd question the choice of pinging a public IP like 8.8.8.8 for a branch SLA. It introduces too many external variables. Your failover is now tied to Google's DNS availability and the routing path to them, which isn't a reliable metric for the health of *your* circuit up to the ISP handoff. You should be probing your ISP's first-hop gateway or a trusted internal target reachable via both paths. Otherwise, you're not just testing your link failure, you're testing the entire internet's stability.



   
ReplyQuote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

Unplugging the cable is a valid hard-failure test. The 1ms SLA trick works, but it's a synthetic condition.

Your bigger gap is the SLA target. Pinging 8.8.8.8 makes your failover dependent on Google and the public internet path. That's not testing your circuit's health, it's testing the internet. Use your ISP's first-hop gateway or an internal target. Otherwise, you'll fail over for problems that aren't yours.


Trust, but verify


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Unplugging the cable proves it fails when there's no cable. Big surprise. Try simulating a flaky ISP router that drops 30% of packets during a video call - that's when your "automatic" failover will probably hesitate and let the call die.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Agree, but the 30% packet loss scenario is what the SLA threshold should catch. The problem is most vendors default to something like 5% loss over 5 minutes. That's useless for real-time apps.

You need to set a tighter loss SLA (e.g., 2% over 30 seconds) and a jitter threshold. Most don't, because it causes flapping. So the failover hesitates.


Numbers don't lie.


   
ReplyQuote
(@cloud_infra_vet)
Honorable Member
Joined: 4 months ago
Posts: 389
 

That `netstat -n` timer check is a solid forensic method, especially for long-lived TCP sessions like database replication or large file transfers. I've used `ss -o` on Linux for the same purpose. The "silent re-establishment" you describe is exactly what happens when the SD-WAN edge device itself maintains state but the underlying IPsec tunnel or ISP circuit renegotiates, causing a new TCP handshake to be proxied.

One nuance I'd add: this check is definitive for TCP, but a lot of modern SaaS traffic is HTTP/3 over QUIC, which is UDP-based. A QUIC connection can migrate paths without the client seeing a reset, so the failover might be truly seamless even if the underlying transport session ID doesn't persist in `netstat`. In those cases, you need to watch the protocol-specific connection ID in client debug logs or a packet capture to see if the flow was truly stateful.



   
ReplyQuote
Page 2 / 2