Skip to content
Notifications
Clear all

Just ran a failover test during business hours. What we learned about user sessions.

10 Posts
9 Users
0 Reactions
33 Views
(@emilyl)
Honorable Member
Joined: 3 months ago
Posts: 527
Topic starter   [#22408]

Hey everyone! I've been lurking for a bit while we set up our new Sophos XGS at our small remote-first company. We're a team of about 25, mostly using Asana and Slack, and I'm the one who got kinda voluntold to help with the IT side of things 😅.

We just did our first planned failover test today, switching from our primary XGS to the secondary unit around 11 AM. I was super nervous about dropping everyone's calls and messing up their work in progress. The actual network cutover was smooth, but we learned something unexpected about user sessions. A bunch of people got logged out of our internal admin tools and even some cloud apps, even though the internet came back right away. Our IT guy said it's because the session states didn't sync perfectly before the failover? I thought High Availability was supposed to be seamless.

So my question for you all is: is this normal? How do you handle failover testing without disrupting everyone's work? Do we need to tweak a specific sync setting for application persistence, or is a brief logout just part of the deal? I want to make our next test less disruptive for the team.

Thx!



   
Quote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

Your IT guy is right about session state sync. Even with HA firewalls, the TCP session table and application-layer persistence data (like which user is tied to which firewall for a specific cloud app) often have a sync delay. The failover is seamless for *network connectivity*, but not for *stateful application sessions*.

A brief logout is often part of the deal if your apps rely on the source IP for session stickiness and that changes. To make it less disruptive, you could look at a few things:
- Check if your XGS has a setting for more aggressive session synchronization.
- Consider offloading session state entirely to an external store, like Redis, for your internal admin tools. That way the firewall's failover doesn't affect it.
- Schedule tests during low-activity periods until you've tuned the sync intervals.

It's a common learning moment - HA keeps the lights on, but true user session resilience usually needs a layered approach beyond the firewall.


sub-100ms or bust


   
ReplyQuote
(@chrisw2)
Reputable Member
Joined: 2 months ago
Posts: 309
 

Yep, that layered approach is key. The "seamless for network, not for state" bit is the trap everyone hits.

We ran into this with Grafana logins. Even with aggressive sync, a session cookie tied to a specific HAProxy backend IP will blow up on failover. Offloading to Redis for our internal stuff was the only real fix.

For cloud apps, you're at the mercy of their session management - sometimes you just have to accept the brief logout as the cost of HA.


Run it yourself.


   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 3 months ago
Posts: 527
Topic starter  

Oh wow, that's really good to know this is a common issue. I'm just starting to learn about our network setup and I would've been totally caught off guard by the logouts, too.

Our team uses Asana heavily for task tracking. Would the failover cause dropped sessions there as well, or is it more about internal tools? I'm wondering if we need to warn people to save their work before a test, even if the internet stays up.



   
ReplyQuote
(@hannahk)
Estimable Member
Joined: 3 months ago
Posts: 173
 

Great question about Asana. In my experience, cloud apps like that are hit-or-miss during a firewall failover. It depends entirely on how *their* servers track your session.

Even though the internet comes right back, the session your team's browser has with Asana was tied to the old firewall's IP or its specific connection path. If Asana's backend sees a sudden change in that, it can invalidate the session for security, thinking it might be a hijack attempt. So yes, I'd absolutely warn people to save any open, unsaved work in their browser tabs before a test.

It's a good habit anyway, and it beats having someone lose a half-written project description!


edge cases matter


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Yep, it's exactly a security feature. The cloud app sees a sudden new source IP for an existing session and kills it. The only way around it is if your HA pair presents a single virtual IP to the outside, and the app never sees the change. Not all setups can do that.

Always tell users to save work before a failover test, even a planned one. The network coming back doesn't mean their sessions will.


Beep boop. Show me the data.


   
ReplyQuote
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
 

Totally normal. The 'seamless' promise is for the network layer, not the app layer. Your IT guy nailed it.

We do our planned failovers on Saturday mornings. Still get a few session drops on some cloud apps, but the disruption is minimal. No way around it unless you can afford a fully stateful sync setup, which is often more complex than it's worth for a small team.

Question for you - did your admin tools and Asana/Slack behave the same way, or was one group worse? Sometimes the internal stuff is easier to fix with better sync settings.


Ask me about hidden egress costs.


   
ReplyQuote
(@emma23)
Reputable Member
Joined: 3 months ago
Posts: 212
 

Yep, totally normal! We see this all the time. Even with good HA, app sessions often drop.

For your next test, I'd recommend sending a quick "save your browser work" warning 10 minutes before. That's become our standard practice. It's less about the network and more about the app security like others said.

You can try tweaking the session sync settings on the XGS, but in my experience, a 2-minute warning saves more headaches than trying to make it 100% seamless for every single app.


Trial first, ask later.


   
ReplyQuote
(@hannahr)
Reputable Member
Joined: 3 months ago
Posts: 285
 

That's a smart schedule. We do ours early Sunday for the same reason.

To answer your question, our internal tools took the hit much worse than the cloud apps. Our old helpdesk system (on-prem) would drop every single login, while services like Slack usually just hiccuped. I think it's because the internal stuff often had tighter, less forgiving session timeouts tied directly to the firewall's connection tracking.

You're right about the complexity for small teams. We looked at a more stateful sync, but the cost and management overhead just weren't justified. Sometimes the simpler fix is adjusting the tool's session timeout to be a bit more forgiving, if you can.


Data is sacred.


   
ReplyQuote
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 441
 

Your point about the internal tools being hit harder really resonates. We've seen the same pattern here.

It often comes down to the design assumptions of the application. A lot of on-prem or internally-hosted tools were built assuming a single, stable connection path, so they tie sessions tightly to specific network attributes like source IP. Cloud-first apps like Slack or Asana are often built for more volatile conditions from the start, like users switching from WiFi to mobile data, so their session management can be a bit more resilient to source changes.

Adjusting the session timeout on the internal tools, if that's an option, is a great practical tip. It doesn't prevent the session from dropping during the failover, but it can give users a bit more grace to re-authenticate without losing their place, which softens the blow considerably.


Stay curious.


   
ReplyQuote