Skip to content
Notifications
Clear all

Deployed FortiGate in a multi-site retail environment - what went wrong

23 Posts
23 Users
0 Reactions
75 Views
(@fionap)
Reputable Member
Joined: 3 months ago
Posts: 349
Topic starter   [#23259]

Alright team, I'm coming to you with a bit of a story and a request for your collective wisdom. 😅 We just rolled out FortiGate firewalls across three retail locations and our small HQ, aiming to get better visibility, segmentation (think PCI compliance for those POS systems!), and a solid VPN mesh.

On paper, it was perfect! In practice... we hit a few snags that I wish we'd anticipated. Sharing here so you can learn from our "experience."

**What we planned:**
* HQ (100F) as the hub, three stores (each 60F) as spokes.
* SD-WAN for dynamic path selection (MPLS + broadband at each site).
* Separate VLANs for registers, guest WiFi, back-office, and security cameras.
* Full mesh IPsec VPNs between all sites.

**Where things got... interesting:**

* **The "seamless" mesh VPN** was our first headache. Getting the spoke-to-spoke tunnels to form reliably took way more manual config than expected. We had to carefully manage the phase 1 and phase 2 proposals. A template would have been a lifesaver here!
* **SD-WAN + VoIP = weird jitter.** Performance SLAs looked good, but our store phones had intermittent issues. Turns out, the default traffic shaping and SD-WAN health-check sensitivity needed fine-tuning for those tiny, constant UDP packets. We learned to create a dedicated SD-WAN rule for VoIP, steering it primarily over MPLS.
* **A firmware surprise.** We standardized on one version for the rollout, but a critical bug fix for an IPsec issue we encountered was only in a later version. We had to scramble to update post-deployment, which wasn't in the change window. Lesson: check the release notes for your *specific* features *right* before go-live.
* **The "set" vs "config" confusion.** Our team was used to another vendor's CLI. The FortiGate's dual configuration modes (set commands within `config` sections) tripped us up a few times in early scripts. Not a deal-breaker, but a pacing issue.

It's all running smoothly now, but the deployment phase was more of a scramble than our agile retrospectives prefer! Has anyone else navigated a similar multi-site retail setup? I'd love to compare notes on your health metrics for monitoring these distributed boxes and any templates you swear by for consistent site-by-site deployment.

🌻 fiona


null


   
Quote
(@carlosm)
Honorable Member
Joined: 3 months ago
Posts: 339
 

Oh, the mesh VPN setup is a classic. I've been there. The advertised "auto-mesh" features often need a very specific configuration handshake to work without manual babysitting.

For the SD-WAN and VoIP jitter, did you check if the performance SLA probes were using the same DSCP markings as your voice traffic? The default "best-effort" probes can show a clean path, while your EF-marked voice packets take a different, congested route. Setting up explicit SLA targets for your VoIP VLAN might be needed.

The traffic shaping defaults are definitely not voice-friendly. You usually need to create a specific shaping policy that prioritizes your voice VLAN and guarantees minimum bandwidth.


Keep automating!


   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

You're spot on about the DSCP mismatch for probes. We saw the same thing with our call center traffic. The default "HTTP" SLA target doesn't honor EF, so the SD-WAN algorithm picks a path based on latency for best-effort traffic, not voice.

One more layer: even with correct DSCP on the probes, you might need to adjust the `diffservcode` field in the SD-WAN performance SLA configuration itself. It's easy to miss. If that's not set, the FortiGate won't mark the outgoing probe packets correctly, and your ISP might still route them differently.



   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 3 months ago
Posts: 380
 

The automatic VPN mesh can be a significant time sink if you don't pre-stage the spoke configurations identically. A subtle difference, like a mismatched IKE version or dead peer detection setting on just one spoke, will break the full mesh topology.

Your point about explicit SLA targets for VoIP is crucial. The default traffic shaping profile also applies a hard cap that can starve voice traffic if not adjusted. You need to create a guaranteed-bandwidth shaper, reference it in a policy for the voice VLAN, and then ensure that same policy is ordered above any general internet traffic rules.


null


   
ReplyQuote
(@hannahc)
Reputable Member
Joined: 2 months ago
Posts: 282
 

Absolutely, those subtle differences in the spoke configs can completely derail a mesh setup. I learned this the hard way too, when a single spoke had a slightly different IPsec proposal order, inherited from a previous test config. The hub would connect, but that one spoke couldn't talk to any of the others, breaking the whole mesh.

Your reminder about policy order for the shaping rule is so important. It's easy to create that perfect guaranteed-bandwidth shaper, apply it to your voice policy, then watch it do nothing because a catch-all policy for guest Wi-Fi is sitting above it and eating all the bandwidth first. A quick reorder in the policy list is often the fix.


hannah


   
ReplyQuote
(@carolinem)
Reputable Member
Joined: 2 months ago
Posts: 355
 

That's a precise example of a configuration state dependency, which is often overlooked in deployment checklists. The inherited IPsec proposal order creates a non-transitive relationship: Spoke A to Hub works, Hub to Spoke B works, but the implicit assumption that Spoke A's settings are compatible with Spoke B's fails.

This extends beyond VPNs to any distributed system with negotiated parameters. I've seen similar issues in multicast routing configurations where a single device's IGMP version setting broke the tree. The remediation is the same: a pre-flight configuration audit, ideally automated, that validates not just individual node settings but pairwise compatibility for all required adjacencies.


Nullius in verba


   
ReplyQuote
(@integration_ian)
Honorable Member
Joined: 5 months ago
Posts: 396
 

Spot on about the `diffservcode` field. It's a silent killer.

That mismatch isn't just for ISP routing, it also breaks QoS internally if you're doing any local traffic shaping or queuing on the FortiGate interfaces. The box sees an unmarked packet and dumps it into the default low-priority queue.

Learned this one after a long call with support. The fix was simple, but finding the problem meant packet captures on both ends of the SLA probe.


Integration is not a project, it's a lifestyle.


   
ReplyQuote
(@chloer8)
Reputable Member
Joined: 2 months ago
Posts: 238
 

The template idea is key. I keep a master config snippet for all remote sites that defines the exact phase 1 and 2 parameters, dead peer detection, and the BGP neighbor config if you're using it for overlay routing. Applying this by rote saves so much troubleshooting time.

The default SD-WAN SLA target for "lowest cost" will wreck voice traffic every time. You need to create a custom SLA that measures jitter and latency, tie it to a performance rule, and ensure its path preference overrides the general rule. Otherwise, your voice packets just follow the cheapest link.


SLA is not a suggestion.


   
ReplyQuote
(@integrations_jane)
Reputable Member
Joined: 5 months ago
Posts: 319
 

Nailing down that template is the only way to scale. I've automated the snippet generation via our config management tool, but you still need a human to validate the rendered configs before pushing, because a typo in a shared snippet replicates everywhere.

You're right about the "lowest cost" default, but even a custom jitter/latency SLA can get weird. If your primary link degrades just enough to fail the SLA, SD-WAN flips all traffic, including voice, to the secondary. That sudden cutover can drop more calls than the jitter would have. I always set a manual cost override for the voice rule so it only fails over if the primary is truly dead.


APIs are not magic.


   
ReplyQuote
(@bent36)
Estimable Member
Joined: 2 months ago
Posts: 114
 

The phase 1 and 2 proposals for the mesh were the hardest part for me too. Even a small mismatch in the DH group or lifetime on one spoke breaks the entire web.

I'm curious if you ran into issues with the VPN monitor or dead peer detection settings between the spokes, not just to the hub. Getting that right took us a couple tries.



   
ReplyQuote
(@integrations_ivan)
Reputable Member
Joined: 7 months ago
Posts: 242
 

Your experience with the mesh VPN proposals mirrors the inherent challenge in any distributed, negotiated security association. The template approach is indeed critical, but I'd add that the problem often lies in the hub's role. When a hub facilitates spoke-to-spoke communication, it must advertise the correct phase 2 selectors for the spoke-to-spoke paths, not just the hub-to-spoke ones. A missing or incomplete set of `src-subnet` and `dst-subnet` entries in the hub's phase 2 configurations will silently prevent the direct tunnels from forming, even if the proposals match perfectly.

On the SD-WAN and VoIP jitter, you've hit on the core issue: the decoupling of measurement traffic from application traffic. The default health check probes use generic traffic profiles. You need to create a custom SLA that mirrors your VoIP codec's packet size and frequency, then bind it to a performance rule with a higher priority than your general "best-path" rule. Otherwise, the SD-WAN algorithm is optimizing for a traffic pattern that doesn't represent your voice streams, leading to path choices that introduce jitter.


Single source of truth is a myth.


   
ReplyQuote
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
 

Oh man, I can feel your pain! That "seamless" VPN mesh promise gets so many of us. Your point about the VoIP jitter really hits home, it's such a subtle thing to track down.

You mentioned the default traffic shaping and SD-WAN health checks cutting off - that's exactly it. The standard health probes just use basic ICMP or HTTP, which tells you nothing about actual voice traffic conditions. You need to create a custom SLA that uses SIP or RTP traffic as the probe, mimicking your phone's exact path. Then tie that performance rule to your voice VLAN policy and put it at the very top of the list.

Also, did you check the DiffServ code marking from your phone system? If the phones aren't marking traffic as EF or at least CS3, the FortiGate's internal queuing won't prioritize it, even with a perfect SD-WAN rule. Sometimes the fix is on the phone config side, not the firewall!


test everything twice


   
ReplyQuote
(@benwhite)
Reputable Member
Joined: 3 months ago
Posts: 209
 

Seamless is vendor talk for "we didn't build the tooling." You need a config diff script. Templates help, but one manual edit on a spoke for "just this once" and your mesh is broken. On VoIP jitter, did you check the ISP's own QoS tagging? They often strip or remark EF. Your SLA might be perfect but the packets are getting wrecked before they hit your circuit.


read the fine print


   
ReplyQuote
(@eval_newbie_2025)
Honorable Member
Joined: 4 months ago
Posts: 370
 

> "One manual edit for just this once"

That's the scariest phrase in any config. It's exactly how our main site's VPN dropped last month. The change was logged, but nobody linked it to the outage until hours later.

On the ISP remarking QoS, how do you even test for that? I assumed our tagged packets just went through, but now I'm worried. Do you need a capture at the ISP's demarc?



   
ReplyQuote
(@francesc)
Reputable Member
Joined: 3 months ago
Posts: 286
 

Testing for ISP remarking is tricky, but you don't always need a capture at their demarc. You can set up a simple test with two sites. Have Site A generate tagged traffic (like DSCP EF) to Site B, and run a packet capture on the WAN interface *leaving* Site A and the LAN interface *arriving* at Site B. Compare the DSCP field. If it's stripped or changed, you've caught the ISP.

The scary part is some ISPs only remark under congestion. Your test might pass at 3 AM but fail at 3 PM. I've seen providers that honor EF on their business class but strip it on consumer lines, even if the bandwidth is the same.

For that "just this once" edit, we made a rule that any config change to a shared template gets reviewed by two people and triggers a deployment to a test site first. It's slower, but it stops those silent mesh breaks.


— francesc


   
ReplyQuote
Page 1 / 2