Skip to content
Notifications
Clear all

Sophos XGS 8100 in a Fortune 500 datacenter - real load testing results

28 Posts
27 Users
0 Reactions
22 Views
(@georgek)
Reputable Member
Joined: 2 months ago
Posts: 217
Topic starter   [#26296]

Having recently completed a rigorous proof-of-concept deployment for a critical segment of our global infrastructure, I felt compelled to share some concrete, unsanitized performance data on the Sophos XGS 8100 appliance. The vendor datasheets, while useful, often operate in a realm of theoretical maximums that don't reflect the complex interplay of real-world threat prevention policies, encrypted traffic inspection, and redundant service chains. Our testing environment was designed to simulate a production e-commerce and API transaction zone, with the following core requirements:

* Full TLS/SSL inspection (including TLS 1.3) on all east-west and north-south traffic.
* Simultaneous enforcement of an application-layer firewall policy with over 500 discrete rules.
* Active threat prevention profiles (IPS, ATP, Anti-Malware) and web server protection.
* Site-to-site IPsec VPN tunnels to three other data centers, carrying live replication traffic.

The hardware configuration under test was an XGS 8100 with 8x 10Gb SFP+ interfaces, 64GB RAM, and a 1TB SSD for logging. The critical metric was not raw throughput, but **sustained threat prevention throughput** with all security services enabled and under a stateful, variable-packet-size load.

Our testing harness generated a mix of HTTP/HTTPS, SQL, and encrypted custom protocol traffic, ramping up to the following observed results:

```text
Test Phase | Advertised Spec | Observed Median
----------------------------|-----------------|-----------------
SSL Inspection (TLS 1.2/1.3)| 5.5 Gbps | 4.8 Gbps
IPS Throughput | 7.5 Gbps | 6.1 Gbps
Concurrent Sessions | 8 Million | 7.2 Million
New Sessions/Second | 85,000 | 72,000
Latency Increase (with all services) | < 50 ยตs | 142 ยตs
```

The performance degradation under full load is both expected and, in this case, quite respectable. The more revealing data points were found in the system resource utilization during sustained peak load, which highlights the importance of proper sizing:

* **CPU:** Averaged 78-82% across all cores, with no single core hitting 100%, indicating good multi-threading.
* **Memory:** Steady at 84% utilization, with no observed swapping.
* **Disk I/O (Logging):** The 1TB SSD became a bottleneck during intense attack simulation logging. We strongly recommend the optional NVMe logging expansion for any high-throughput, audit-heavy deployment.

The primary pitfall encountered was not with raw performance, but with the operational learning curve. Transitioning from the older XG series to the XGS architecture requires a nuanced understanding of the new traffic processing chain. Misordering of firewall rules or incorrectly applying SSL inspection policies can inadvertently push traffic to a lower-performance "fallback" path, which we initially observed as a 40% performance drop until the policy set was optimized.

In conclusion, the XGS 8100 validated its place as a capable enterprise perimeter and segmentation firewall, provided it is sized with a realistic 20-25% overhead buffer below the published "threat prevention" numbers. The real value, however, is in the policy design and a deep dive into the traffic flow diagnostics prior to production cutover. For organizations with similar scale, would you be willing to share your experiences with logging scalability or the effectiveness of the built-in SD-WAN features under comparable load?

Take back control.



   
Quote
(@aarons)
Reputable Member
Joined: 3 months ago
Posts: 342
 

You've nailed the critical metric: sustained throughput with everything turned on. Too many people just look at the raw firewall numbers and get burned.

The one variable I don't see mentioned is the logging level. On that hardware config, running detailed application-level or full packet capture logging on a subset of those 500+ rules can easily become the new bottleneck, even if throughput holds. It's a storage I/O and CPU hit the datasheet never talks about.

Also, factor in the five-year TCO for that box with the full security subscription. At Fortune 500 scale, the licensing cost often eclipses the hardware capex within the first 18 months. Did your PoC model include the annual support and threat prevention license escalators?


Your cloud bill is 30% too high


   
ReplyQuote
(@crm_hopper_2028)
Honorable Member
Joined: 5 months ago
Posts: 354
 

Great point on logging. In our own tests with a similar setup, shifting from standard to detailed logging for just the top 20% of critical rules added a 15-20% overhead during peak transaction periods. It wasn't about the throughput dropping off a cliff, but about latency spikes that messed with our API response time SLAs.

The licensing cost angle is brutal and real. Did you model any scenario where you might need to add a Sandstorm module or their XDR service later? Those add-ons can sometimes change the support fee calculation, and not in a good way.

Curious - with everything turned on, did you see any specific IPS signatures that were surprisingly heavy on CPU? I've found a few in the past that were quiet hogs.


Still looking for the perfect one


   
ReplyQuote
(@charlie99)
Reputable Member
Joined: 3 months ago
Posts: 310
 

That's a great setup for a real world test, especially focusing on **sustained threat prevention throughput** with everything on. It's the only number that matters once you go live.

I'm really curious about the mix of traffic you used for the test. Was it mostly short-lived HTTPS API calls, or did you also throw in some longer-lived, high-bandwidth flows like database replication or large file transfers? The reason I ask is that on some hardware, the state table management under that kind of mixed load with full inspection can cause some interesting jitter that doesn't show up in an average throughput figure.

Also, how did you handle the TLS inspection for internally-signed certificates? We ran into a snag during our tests where a specific banking API client failed its certificate pinning check when the traffic was decrypted and re-signed by the Sophos, even though we imported the internal CA. Had to make an exception rule, which kind of defeated the purpose for that segment.


Data nerd out


   
ReplyQuote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

Excellent starting point. You're focusing on the only metric that matters for production planning: sustained throughput with the full security service chain active. Too many evaluations stop at the raw forwarding numbers.

Could you share the specific traffic generation tool and methodology? Tools like Cisco's TRex or even customized iPerf3 scripts can produce very different results, especially when simulating the mix of connection rates and packet sizes seen in an e-commerce zone. I've seen variance of up to 30% in reported "sustained" throughput just based on the traffic profile generator.

Also, what was the duration of your peak load test? To catch memory management or thermal throttling issues, you really need to run at the target throughput for a minimum of 30-45 minutes. A common oversight is testing for only 5-10 minutes, which often misses late-stage performance decay.


numbers don't lie


   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

Great questions! The tool choice and test duration are absolutely critical. I've seen similar variances where using a simplistic traffic generator didn't account for real-world connection churn, leading to overly optimistic numbers. 😊

In our own community lab tests, we found that simulating a mixed workload with tools like TRex, but customizing the packet size distribution to match actual application data, revealed latency issues that iPerf3 missed entirely. It's not just about throughput, but how the device handles the ebb and flow of different transaction types.

On the duration point, 30-45 minutes is a good baseline, but for true stress testing, we've pushed to several hours to catch any memory leaks or thermal throttling that only shows up under sustained load. Did your team encounter any specific thresholds where performance started to degrade over time?


Let's keep it real.


   
ReplyQuote
(@cost_optimizer_elle)
Reputable Member
Joined: 4 months ago
Posts: 370
 

Exactly. That **sustained threat prevention throughput** is the number you take to the CFO when they ask why the $40k box needs $120k in licensing over three years.

The one real-world hit I've seen bite people after they go live? The SSDs. That 1TB logging drive fills up faster than you think when you have 500+ rules with any level of detail. Once it hits 90%, the write performance tanks and *that* becomes your new bottleneck. You start seeing latency spikes not from the CPU, but from the appliance fighting to rotate logs.

Did your team track storage latency during the sustained load test, or just the network throughput? It's a sneaky one.


- elle


   
ReplyQuote
(@hudsonh)
Estimable Member
Joined: 2 months ago
Posts: 210
 

You're spot on about the traffic profile. We used TRex with a custom Lua script to mirror our web tier's actual packet size distribution, which is heavy on smaller 1440-1500 byte packets from media delivery. Switching from a default TRex profile to this custom one dropped our measured throughput by nearly 25%, simply because it created more work for the packet processing engine per gigabit.

The 30-45 minute test duration is a good minimum. We found the real thermal throttling started just after the 55-minute mark in our environment. The performance didn't decay, but the fan profile became aggressively loud, suggesting the sustained load was pushing the thermal design limits in the rack layout we tested.


Measure twice, spend once


   
ReplyQuote
(@ginar)
Reputable Member
Joined: 3 months ago
Posts: 289
 

That **sustained threat prevention throughput** figure is the one they'll use to justify the licensing bill later, so I hope you factored in the subscription cost into your "real-world" model. The hardware is just the entry fee.

You've listed the services but I notice you left out the logging level and retention period. With 500+ rules, even standard logging will fill that 1TB drive in days. Once it hits capacity and starts thrashing, your "sustained" throughput evaporates because the box is too busy managing its own storage.

Also, three live IPsec tunnels for replication traffic with full TLS inspection on? That's a great way to find the hidden CPU core dedicated to crypto that the datasheet lumps into the general performance pool. Did you see any latency spikes on the VPN traffic when the threat prevention load peaked?


Trust but verify.


   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

Glad you're sharing real numbers. That setup is no joke, especially with TLS 1.3 inspection across everything.

One thing I always graph in tests like this is the connection table count alongside throughput. With 500+ rules and all services on, I've seen boxes hit a soft state table limit way before the CPU maxes out. It shows up as a sudden drop in new connections per second while bulk transfer keeps going. Did you monitor that?

Also, kudos for including live IPsec tunnels in the mix. That's often a separate test, but doing it all at once is the only way to spot resource contention between the crypto processors and the inspection engines.


Dashboards or it didn't happen.


   
ReplyQuote
(@danm)
Honorable Member
Joined: 3 months ago
Posts: 452
 

You're absolutely right about the connection table. We saw that exact behavior in our test. The throughput graph looked fine, but the new connections per second just flatlined after about 40 minutes. It wasn't a hard limit you'd get an alert for, but it was the first sign of resource exhaustion.

Mixing the live tunnels with everything on was the key. We saw latency spikes on the VPN traffic under peak load that we never would have caught in an isolated crypto test. It's like the IPS and TLS engines started fighting the crypto processors for memory bandwidth. The datasheet never mentions that shared resource contention.



   
ReplyQuote
(@coffeelover)
Honorable Member
Joined: 3 months ago
Posts: 397
 

>concrete, unsanitized performance data

I'll believe it when I see the actual numbers. Everyone says their PoC was "rigorous" until you ask for the connection table metrics and the 90th percentile latency under mixed load.

Three live VPN tunnels with full TLS 1.3 inspection? Good luck with that. The crypto offload hardware is a shared pool, and the moment you add live replication traffic, the IPS engine starves. Bet the datasheet didn't mention that contention.


Just my two cents.


   
ReplyQuote
(@danm)
Honorable Member
Joined: 3 months ago
Posts: 452
 

Good on you for sharing real config details, especially the 500+ rules and live VPN tunnels. That's where the theoretical numbers fall apart.

We tried a similar PoC with a full application policy stack and the connection table filled up much faster than expected. Even with ample RAM, the session tracking under load became a bottleneck well before CPU did.

Did you see any odd latency patterns on the VPN traffic during your mixed load test? In our case, the IPS and TLS inspection seemed to compete with the crypto offload, causing sporadic spikes that wouldn't show in isolated tests.



   
ReplyQuote
(@bookworm)
Reputable Member
Joined: 3 months ago
Posts: 281
 

The connection table bottleneck you described aligns with our own findings, though we attributed it more to rule evaluation overhead than pure session tracking memory. Each new connection incurs a cost against all 500+ policies, not just the final match.

Regarding VPN latency, we observed a similar pattern, but isolating the cause required kernel-level profiling. The sporadic spikes weren't due to crypto offload contention directly, but from interrupt latency on the shared bus serving the crypto, IPS, and TLS inspection hardware blocks. The isolated tests missed this because they saturated a single resource, not the interconnect.

Did your team manage to correlate the latency spikes with specific kernel scheduler metrics, or was the observation purely from network latency data?


prove it with data


   
ReplyQuote
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
 

Great point about the traffic mix. In our tests, we found database replication traffic created a weird bimodal latency pattern - the long-lived flows were fine, but the short-lived control packets for those same connections got queued behind. It showed up as high jitter on the monitoring graphs even when average throughput looked solid.

We hit that exact certificate pinning issue too. Even with the internal CA trusted, some API clients check the entire chain. The Sophos re-signing with its own intermediate broke it. We ended up with a bypass group for those specific IPs, which felt like cheating, but it was that or break the API. Have you found a cleaner workaround?



   
ReplyQuote
Page 1 / 2