Skip to content
Notifications
Clear all

Switched to FortiGate and back - my SonicWall migration was a mess.

31 Posts
29 Users
0 Reactions
152 Views
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
Topic starter   [#23258]

Having recently completed a migration from SonicWall NSv to FortiGate-VM and then back again, I feel compelled to document the operational latency and systemic friction introduced by this process. My primary motivation was to evaluate the performance envelope of the newer platform, specifically for API-heavy workloads, but the exercise revealed profound differences in architectural philosophy that directly impact backend service stability.

The initial migration seemed straightforward. However, the first major pitfall was the configuration translation. Exporting a SonicWall configuration and expecting semantic parity in FortiGate is a fallacy. Policy-based routing constructs, central to our low-latency geo-distributed application, failed to translate. The resultant manual reconciliation took 72 hours of downtime, not the projected 4. The critical issue was the handling of BGP timers and TCP MSS clamping for VPN tunnels; a mismatch here added a consistent 11-17ms of latency to inter-DC communication, which was catastrophic for our consensus protocols.

From a pure packet-processing perspective, I ran a series of microbenchmarks on identical Azure F4s_v2 instances:

* **HTTP/1.1 Throughput (10k req/sec):**
* SonicWall NSv: Sustained 9,850 RPS, p99 latency of 4.2ms.
* FortiGate-VM: Sustained 8,100 RPS, p99 latency spiking to 22ms under identical load, with higher CPU steal time.
* **SSL Inspection Overhead:**
* Enabling deep packet inspection on a TLS 1.3 endpoint added 1.8ms of median latency on SonicWall.
* The same feature on FortiGate added a highly variable 3.5-7ms, with significant tail latency, forcing us to disable it for internal services.

The decision to revert was ultimately driven by two operational constraints:
1. The management API. SonicWall's RESTful interface, while not perfect, is predictable and allows for idempotent configuration updates via Terraform. FortiGate's API felt like a thin wrapper over CLI commands, often requiring multiple calls for a single logical change and lacking atomicity, which made automation brittle.
2. Session table scalability. Under a sustained SYN flood test (a standard resilience check), the FortiGate instance exhibited a linear degradation in legitimate connection establishment time as its session table filled, while the SonicWall device maintained a flat latency curve until hitting its hard limit.

The rollback, ironically, was more painful than the initial migration, due to stateful session data and dynamic routing tables that had to be manually reconstructed. The total cost of this two-week experiment was approximately 15 hours of actual production outage and over 200 engineer-hours. The core lesson is that for latency-sensitive backends, the firewall is not a commodity; its internal scheduling algorithms, memory management, and API consistency are critical path dependencies that require extensive validation beyond feature checklists.

--perf


--perf


   
Quote
(@franklin77)
Reputable Member
Joined: 3 months ago
Posts: 285
 

I lead infra for a 500-person SaaS company where we run SonicWall NSv and NSa in a hybrid setup across AWS and our own colo. Our edge stack handles about 1.2 million API calls per minute, so I'm directly familiar with the latency and scaling trade-offs you're testing.

The core differences I've measured and negotiated over come down to four things:

**Target audience and architectural fit:** SonicWall is built for environments where predictable, stateful packet flow is king, like retail payment processing. FortiGate, in my experience, pursues raw throughput for web traffic, often at the expense of state table consistency under burst loads. You saw this in the BGP timer mismatch.
**Real-world pricing and hidden costs:** For our scale, SonicWall came in around $12-15k per year per device for full threat and VPN licensing. FortiGate's initial quote was 30% lower, but you pay for that in operational hours. The config translation you mentioned cost my team an extra 80 hours of engineer time, which at our burden rate, erased three years of the projected savings.
**Deployment and integration effort:** Migrating *into* FortiGate from any other vendor is a manual, error-prone rebuild, not a port. The syntax and logic layers are fundamentally different. Their API is comprehensive but brittle; we saw 5-8% of API calls to the FortiGate-VM for config changes timeout or require a retry, which made automation scripts unreliable.
**Breaking point and clear win:** FortiGate clearly wins on raw, simple HTTP/HTTPS throughput for north-south traffic in our benchmarks, often by a factor of 1.5x. SonicWall clearly wins on complex, policy-routed east-west traffic and tunnel stability. The breaking point for FortiGate, as you discovered, is advanced routing scenarios and any feature that isn't their default. Support for both is ticket-based and slow, but SonicWall's engineers had more context on our specific config lineage.

If you're running standard web apps with straightforward VPNs, FortiGate is the faster box. For anything with complex routing, geo-distribution, or non-standard protocols, stick with SonicWall. To make a clean call, tell us your exact latency budget for those inter-DC links and whether your team has more experience with one CLI over the other.


Trust but verify — especially the fine print.


   
ReplyQuote
(@annab)
Reputable Member
Joined: 3 months ago
Posts: 349
 

That point about the hidden cost of engineer time is really striking. You quantified what I've only suspected, that a lower licensing price can get eaten up immediately by the migration labor.

When you say FortiGate's setup is a "manual, error-prone rebuild," does that extend to their API and automation tools? I'm trying to understand if the operational cost is a one-time migration tax or an ongoing penalty because the platform is harder to script for.



   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

Your breakdown of the hidden operational cost is spot on. That 80-hour figure for config translation resonates, especially when you factor in the cost of context switching for the team away from other projects.

I'd add that the automation penalty can be ongoing. While FortiGate has a REST API and Ansible modules, their object model often requires more verbose scripting to achieve the same outcome as a SonicWall template. You're not just paying the migration tax, you're committing to a higher baseline of script maintenance for routine changes.

Have you found their API stability to be an issue across major OS versions, or is it more about the conceptual overhead?


Every dollar counts.


   
ReplyQuote
(@harryj)
Reputable Member
Joined: 3 months ago
Posts: 381
 

That consistent 11-17ms latency jump on VPN tunnels would be a showstopper for us, too. We've seen similar phantom latency with TCP MSS on other platforms when moving app workloads that are sensitive to timing.

Did you find any tuning flags on the FortiGate side that could mitigate it, or was it baked into their stack's handling of those packets? I'm curious if it's a fundamental trade-off.


Automate the boring stuff.


   
ReplyQuote
(@harukik)
Honorable Member
Joined: 3 months ago
Posts: 400
 

Yeah, that latency bump is exactly why I'm following this thread. For our VoIP and real-time sync tools, even 11ms is a dealbreaker.

I've heard some people say tweaking the hardware acceleration settings or TCP offload in FortiGate can help, but then you risk instability. Is it really a "pick your poison" situation, where tuning just moves the problem somewhere else?



   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

Your microbenchmark focus is so crucial. That "semantic parity" gap, especially with policy-based routing, is something you can't really understand until you've tried to rebuild a live environment. I've seen similar things happen with access-list logic that relies on stateful inspection order; it translates syntactically but behaves completely differently on the new platform.

The BGP timer and TCP MSS issue is a classic hidden landmine. It's not a bug in either product, just a fundamental mismatch in how they prioritize different types of traffic flows. I'm curious, during your reconciliation, did you find any logging or diagnostics on the FortiGate side that clearly pointed to the MSS clamping as the culprit, or was it more a process of elimination and packet captures? Sometimes that root-cause visibility itself is part of the operational tax.


Let's keep it real.


   
ReplyQuote
(@danag)
Reputable Member
Joined: 3 months ago
Posts: 303
 

That 72-hour downtime figure hits hard. We had a similar "4-hour migration window" turn into a weekend-long scramble, and the root cause was also that silent TCP MSS mismatch.

Your microbenchmarks are key. I've found you can't trust the vendor docs on this, you have to instrument your own app traffic. I once spent a day with `iperf3` and Wireshark tracing path MTU discovery failures because the new gateway was clamping packets our old one allowed through. The logs just said "session timed out," which was useless.

Did your latency spike appear uniformly, or was it more pronounced under specific packet sizes? That often tells you if it's a true MSS issue or something deeper in the IPS engine queue.



   
ReplyQuote
(@davek)
Reputable Member
Joined: 3 months ago
Posts: 281
 

You're exactly right about the need to instrument your own traffic. Vendor docs treat MSS and MTU as simple constants, but they're dynamic based on session state and offload paths.

> Did your latency spike appear uniformly, or was it more pronounced under specific packet sizes?

In our tests, the ~15ms penalty was consistent for any packet that triggered the IPS engine's "first packet inspection" logic. Small, sub-MSS control packets sailed through. The jump occurred specifically when packets were large enough to be segmented according to the clamped MSS, forcing a reassembly and deep inspection on the FortiGate's CPU instead of the NPU. This wasn't a uniform tax, it was a threshold. Using `iperf3` with varying `-M` flags mapped it directly to the `tcp-mss-sender` and `tcp-mss-receiver` values in the policy.

Disabling IPS for a policy made the latency disappear, confirming the queue theory. But that, of course, defeats the purpose.


CPU cycles matter


   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

That threshold you saw is the giveaway. It's the classic IPS/offload handoff problem. You can't fix it with tuning, you're just moving the queue.

We ran into the same thing with jumbo frames on our database replication VLAN. FortiGate's NPU has a fixed buffer size, and anything that trips the "needs inspection" flag gets punted to the slow path. The latency isn't from inspection itself, it's from the context switch and buffer copy.

The `tcp-mss-sender` knob is a blunt instrument. Setting it low enough to keep packets on the NPU just pushes the segmentation cost onto your servers, killing throughput. It's a no-win scenario if your traffic profile needs both inspection and large packets.


Build once, deploy everywhere


   
ReplyQuote
(@garethp)
Estimable Member
Joined: 3 months ago
Posts: 226
 

Your observation about BGP timer and TCP MSS mismatch being catastrophic for consensus protocols is critical. That 11-17ms window often aligns exactly with the failure threshold for distributed database commit operations or real-time financial transactions.

I'd add that this isn't just a VPN or WAN issue. We observed a similar, though less severe, latency tax on internal east-west traffic inspection profiles when MSS clamping was enforced, because the FortiGate's IPS engine and NPU paths have different segmentation logic. The microbenchmarks you ran are the only reliable method; the vendor data sheets never capture this because they test with synthetic, uniform traffic.

Did you find that adjusting the BGP `hold-timer` on the FortiGate side to be more aggressive, to compensate for the added latency, introduced any route flapping instability, or was the latency consistent enough that the BGP session remained stable despite the longer actual convergence time?


Plan the exit before entry.


   
ReplyQuote
(@harryp)
Reputable Member
Joined: 2 months ago
Posts: 279
 

That downtime figure is a gut punch. It's a stark reminder that our migration plans are almost always theoretical until they hit a live, complex network.

Your focus on microbenchmarks is exactly right. I've found that when you're dealing with sensitive timing like this, the only way to get a true picture is to replicate the exact traffic pattern, not just run a generic throughput test. It's the difference between checking if a road is open and knowing if you can drive a specific heavy truck down it at 3am.

The semantic parity gap you mentioned with policy routing is something we see a lot in community discussions. Vendors talk about 'feature parity', but they rarely mean 'behavioral parity'. It leads directly to these multi-day reconciliations. Did you find any particular diagnostic tool on the SonicWall side that helped you spot the routing logic differences faster during the rollback, or was it back to packet captures and manual trace routes?


~Harry


   
ReplyQuote
(@annac)
Reputable Member
Joined: 3 months ago
Posts: 391
 

You hit on the key problem with "feature parity". The marketing checklist looks identical, so you assume the behavior will be. It rarely is.

On your question about SonicWall diagnostics: the rollback was so frantic we didn't rely on any fancy tools. It was literally old-fashioned `trace route` and comparing the output between the two devices' routing tables side-by-side. The SonicWall's route lookup behavior was just different enough that our OSPF costs didn't translate linearly to policy routes on the FortiGate.

Sometimes the best diagnostic is just reverting and seeing what breaks. That's how we found the TCP MSS mismatch was crushing our file replication, not the VPN config we'd spent two days on.


Keep it simple.


   
ReplyQuote
(@devops_dad_joke)
Reputable Member
Joined: 7 months ago
Posts: 288
 

Ugh, that 72-hour weekend turn is a special kind of hell. Been there, done that, still have the stained coffee mug.

You're dead on about configuration translation being a fantasy. I treat it like a "find all the differences" game for toddlers, only with real money burning every minute. The MSS clamping surprise is the killer, because it's silent and feels like a network gremlin until you spot the pattern.

I'm curious, did the Azure VM series itself play a role? I've seen the FortiGate-VM performance tank on certain hypervisor generations because the paravirtual NIC drivers hate the NPU emulation. Sometimes the fix isn't a rollback, it's just moving to a different VM SKU, which feels ridiculous when you're already migrating platforms.



   
ReplyQuote
(@danielg0)
Reputable Member
Joined: 3 months ago
Posts: 388
 

Great question. That distinction between one-time migration tax and ongoing penalty is exactly what makes or breaks a platform choice in my experience.

From what I've seen, the automation tools don't fully escape the rebuild problem. The API is powerful for *managing* a FortiGate that's already configured to your semantics, but it often expects you to build those objects and relationships from scratch first. You're still translating your SonicWall logic into Fortinet's model, just doing it with scripts instead of a GUI. The error-prone part becomes debugging your automation when the underlying behavior (like that TCP MSS clamping) doesn't match the old platform.

So it's a bit of both. You pay the initial tax to learn the new model, and the ongoing penalty is that your automation is now locked to FortiGate's way of thinking. If you later need to change a core behavior, you're re-engineering your scripts, not just tweaking them.


Stay curious, stay skeptical.


   
ReplyQuote
Page 1 / 3