Skip to content
Notifications
Clear all

Just ran a failover test during business hours. What we learned about user sessions.

2 Posts
2 Users
0 Reactions
1 Views
(@emilyl)
Reputable Member
Joined: 2 weeks ago
Posts: 148
Topic starter   [#22408]

Hey everyone! I've been lurking for a bit while we set up our new Sophos XGS at our small remote-first company. We're a team of about 25, mostly using Asana and Slack, and I'm the one who got kinda voluntold to help with the IT side of things 😅.

We just did our first planned failover test today, switching from our primary XGS to the secondary unit around 11 AM. I was super nervous about dropping everyone's calls and messing up their work in progress. The actual network cutover was smooth, but we learned something unexpected about user sessions. A bunch of people got logged out of our internal admin tools and even some cloud apps, even though the internet came back right away. Our IT guy said it's because the session states didn't sync perfectly before the failover? I thought High Availability was supposed to be seamless.

So my question for you all is: is this normal? How do you handle failover testing without disrupting everyone's work? Do we need to tweak a specific sync setting for application persistence, or is a brief logout just part of the deal? I want to make our next test less disruptive for the team.

Thx!



   
Quote
(@backend_latency_queen)
Reputable Member
Joined: 2 months ago
Posts: 205
 

Your IT guy is right about session state sync. Even with HA firewalls, the TCP session table and application-layer persistence data (like which user is tied to which firewall for a specific cloud app) often have a sync delay. The failover is seamless for *network connectivity*, but not for *stateful application sessions*.

A brief logout is often part of the deal if your apps rely on the source IP for session stickiness and that changes. To make it less disruptive, you could look at a few things:
- Check if your XGS has a setting for more aggressive session synchronization.
- Consider offloading session state entirely to an external store, like Redis, for your internal admin tools. That way the firewall's failover doesn't affect it.
- Schedule tests during low-activity periods until you've tuned the sync intervals.

It's a common learning moment - HA keeps the lights on, but true user session resilience usually needs a layered approach beyond the firewall.


sub-100ms or bust


   
ReplyQuote