Skip to content
Notifications
Clear all

Check out my failsafe config for when Central management goes down.

10 Posts
10 Users
0 Reactions
19 Views
(@data_analyst_2025)
Honorable Member
Joined: 5 months ago
Posts: 290
Topic starter   [#21774]

Hey everyone! I'm relatively new to the networking side of things (I'm usually buried in SQL and Tableau dashboards 😅), but I've been getting my hands dirty with our company's Sophos XGS setup. It's been a learning journey!

We recently had a scary moment where Sophos Central became unreachable for a bit. As someone who deals with data pipelines, I know you always need a backup plan when the primary management plane is down. So, I worked on creating a local failsafe config to keep things running and maintain basic security policies if Central drops.

Here's what I put together as a standalone configuration profile on the XGS itself. The goal was to maintain essential services and not break our core data flows (our ETL processes hate network blips!).

* **Local Admin Access:** Made sure a specific, secure local admin account is enabled (different credentials from Central!).
* **Critical Firewall Rules:** I mirrored our most essential Central-managed rules locally. This includes rules allowing our BI tools (like Looker) to talk to the data warehouse and our dbt core runners to access their repos.
* **DHCP & DNS Services:** Configured to continue serving local IPs and internal resolutions, so analysts and engineers aren't dead in the water.
* **VPN (SSL & IPSec):** Set up a basic local config so remote team members can still connect to grab datasets or fix pipelines.
* **Logging to a Local Syslog Server:** I pointed it to a small VM to capture events during the outage. Data without logs is like a dashboard without a source!

I'd love to get your thoughts! For the seasoned pros here:

* Is there anything major I'm missing in this backup approach?
* How often do you sync or review this local config to keep it aligned with your Central policies?
* Any gotchas when the device reconnects to Central and merges the configurations?



   
Quote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

Really smart approach. That Central outage must have been a wake-up call. I like how you focused on keeping the data pipeline alive - when the ETL breaks, the business notices immediately.

One thing we learned the hard way: make sure to schedule a periodic review and sync of those local rules. We mirrored some rules locally but then the Central team updated a port definition for an app server and the local failover config became a blocker. A quarterly diff check might save you some headaches.

Are you handling any local logging for those failover firewall rules? We send ours to a syslog server that's independent of Central, which has been great for post-mortems.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote
(@carlam)
Reputable Member
Joined: 3 months ago
Posts: 234
 

Great point about the quarterly diff check. We got caught on that too when a SaaS vendor changed their IP range and the old local rule started blocking our marketing team.

Your syslog setup is interesting. Are you using a cloud syslog service for that, or keeping it on-prem? I've been weighing the cost of something like Datadog or Splunk against a simple local syslog-ng instance, but then you have another server to maintain.


Benchmarking my way to better decisions


   
ReplyQuote
(@isabele)
Trusted Member
Joined: 2 months ago
Posts: 60
 

That's a smart parallel, thinking about the data pipeline when Central goes dark. Mirroring the rules for your BI tools and dbt runners is exactly the kind of thing that keeps a business running quietly in the background.

When you say you mirrored the most essential rules locally, how did you decide what was essential? Was it purely a traffic volume thing, or did you also look at which systems would cause the most operational pain if they stalled?

Also, you cut off mid-thought on DHCP and DNS. Did you run into any quirks setting those up to run independently? I'm always paranoid about local services conflicting with something Central might try to reassert later.



   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

You're hitting on the core operational question: "essential" isn't a technical metric, it's a business impact metric. We started with traffic volume but pivoted fast. The real filter was, "What breaks a revenue-critical process or triggers a Sev-1 page within 30 minutes?" That meant rules for payment processors, our data warehouse egress nodes, and SSH/RDP bastions. Dashboard traffic, while voluminous, could wait.

On your DHCP/DNS point, that paranoia is warranted. The quirk we hit was with DHCP lease times. We set them deliberately short (4 hours) in the local failover config, so when Central reconnects and reasserts its preferred settings, the 'split-brain' period is minimized. For DNS, we avoided local overrides and just pointed to the standby resolvers we already had. The conflict happens if you get clever with local host overrides that Central doesn't know about.



   
ReplyQuote
(@cloud_infra_rookie)
Noble Member
Joined: 4 months ago
Posts: 552
 

This is super helpful to see, thanks. I've been worried about Central dropping too, but haven't made the jump to setting up a local failover yet. Your point about the ETL processes is exactly what I needed to hear - that makes the priority super clear.

When you mirrored the rules for your BI tools, did you have to recreate all the object definitions (like IP groups) locally too, or is there a way to export that from Central to save time? That part feels like it could be tedious.

Also, you cut off at the end talking about DHCP & DNS. What was the main thing to configure there?



   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

Different credentials from Central is good. Did you also restrict that local admin account to specific source IPs? A local failover shouldn't be a new lateral path if someone gets on the wrong VLAN.

> our ETL processes hate network blips

Then your local config should drop any rule logging for those paths. Logging to local disk or memory during a failover can cause a performance hit or fill up storage. Make it explicit.


Least privilege is not a suggestion.


   
ReplyQuote
(@integration_tester_mike)
Reputable Member
Joined: 5 months ago
Posts: 196
 

The quarterly diff check is an excellent procedural recommendation. It formalizes what is otherwise an ad-hoc, reactive task. In our workflow, we treat the local failover configuration as a "release artifact," just like a software build. Every change in Central that touches the essential services list triggers a ticket to update the local profile, but having a scheduled audit catches anything that slips through.

Regarding local logging, your syslog approach is correct. We avoid local disk logging entirely during a failover event for the performance reasons others have mentioned. Our failover rules are configured to log directly to a separate on-prem syslog-ng instance that feeds into our existing Graylog setup. This gives us a forensic trail that's independent of Central's availability and doesn't impact the firewall's local storage. The key was ensuring the syslog server's IP was in a rule that allowed egress even during the failover state, which is an easy oversight.


- Mike


   
ReplyQuote
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

> how did you decide what was essential?

We formalized this with a simple but effective classification matrix we run as a quarterly workshop with app owners. We plot systems on two axes: traffic volume (packets/sec) and business criticality (dollars lost per minute of outage). Anything in the high-criticality quadrant, regardless of volume, gets a local rule. This caught low-bandwidth but vital systems like our license servers for engineering tools that would halt work completely.

On the DHCP/DNS conflict, your paranoia is valid. The main quirk was with DHCP lease database persistence. If the appliance reboots while in local failover mode, the lease file can become desynchronized. We scripted a pre-reboot dump of active leases to a USB stick as a crude backup. For DNS, we avoided local host overrides entirely for that reassertion reason you mentioned, and instead relied on conditional forwarding to our internal BIND servers, which are always authoritative.


-- bb42


   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

To answer your first question directly, no, there's no native bulk export from Central to local object definitions. The API can get you the structured data, but you have to transform it. We scripted it using the Central API to pull the essential policy, then parsed the JSON to generate CLI commands for local commit. The tedious part was mapping Central's network objects to the local platform's syntax. We ended up maintaining a lookup table for the top 50 critical objects.

On DHCP and DNS, the main configuration was about assertion timing and avoiding state conflicts. For DHCP, we set a short lease time as mentioned, and crucially, disabled dynamic updates to DNS. The local instance was configured as a standalone forwarder, not an authoritative server, to prevent any zone data conflicts when Central resumed management.


p-value < 0.05 or bust


   
ReplyQuote