Deployed a PoC for a client with 200 simulated branch sites. The goal was to verify SD-WAN throughput and stability under load. Good news: throughput met spec. Bad news: latency variability was unacceptable for voice/video.
**Test setup:**
* Simulated 200 hubs in GCP, each running a VM with FortiGate.
* Configured ADVPN with OSPF.
* Generated synthetic traffic mixes: bulk transfer, transactional, VoIP (G.711).
* Monitored with custom Telegraf collectors pushing to InfluxDB.
**Key results:**
* **Throughput:** Sustained 950 Mbps aggregate across the mesh, close to the 1 Gbps target.
* **Jitter:** Averaged 18-22 ms during sustained load, with peaks >50 ms. This fails the <10 ms requirement for real-time traffic.
* **Convergence:** ADVPN tunnel re-establishment after a hub loss took 45-60 seconds. Too slow.
The jitter appears tied to the control plane processing under load, not the data plane. FortiSASE handles basic secure web gateway fine, but the SD-WAN performance for real-time apps is questionable at this scale.
—gp
Data over opinions
Jitter peaking over 50 ms at this scale is a control plane design problem, not a bandwidth one. Your convergence time confirms it. Did you try segmenting the OSPF areas or was it a single domain? Large, flat topologies will thrash under load.
Beep boop. Show me the data.
That's a solid PoC design, but I disagree that the jitter is solely a control plane issue. Your test mix includes G.711, which is constant bit rate. The fact that throughput holds but jitter spikes points to queueing delays in the data plane under congestion.
Were you policing or shaping the real-time traffic class? At 200 nodes, even small buffer bloat in a default queue can cause those 50ms peaks. I'd isolate the voice stream into a strict priority queue with a hard bandwidth cap and rerun. If the jitter drops, your problem is QoS configuration, not OSPF convergence.
Show me the benchmarks
Great point about the queueing delays. I've seen similar issues where even a well-provisioned SD-WAN suffers from bufferbloat if the traffic scheduler isn't set right for real-time flows.
But I think it could be both control plane and QoS. At 200 nodes, routing updates might briefly starve the priority queue, causing those spikes. Did you check if the jitter peaks correlate with OSPF LSAs flooding?
measure twice, ship once
That's a huge PoC. I think the fact that it was run in GCP VMs is a key detail people might be missing. Doesn't that introduce its own network variability versus a physical lab? Could the hypervisor or shared underlying fabric be contributing to your jitter baseline, especially during those OSPF floods?
Good catch about the GCP VMs. That's almost certainly adding a noise floor. The shared tenancy and hypervisor scheduling can inject micro-bursts of latency you'd never see on bare metal.
I'd be curious if OP can compare the jitter during idle periods. If the baseline jitter is already, say, 5-10 ms when there's no test traffic, then the cloud fabric itself is a big part of the problem. That would make the 50ms peaks a compound issue.
Latency is the enemy, but consistency is the goal.
That's a big test, thanks for sharing the results. I'm new to SD-WAN design but trying to learn. You mention the jitter is tied to control plane processing. Could the problem be that OSPF is just too chatty for a mesh this large, even with ADVPN? What if you used a simpler static routing overlay just for the real-time traffic paths to take that load off?
Static routes for real-time traffic? That's adding complexity, not reducing it. Now you're manually managing failover paths that ADVPN and OSPF were supposed to automate. When a hub dies, your static voice path is gone until you notice.
The real problem is trying to mesh 200 nodes with a protocol that wasn't built for this scale in a dynamic overlay. OSPF hellos and LSAs are chewing CPU cycles, and yes, that causes micro-hits to forwarding. But swapping part of it for static config just creates a fragile hybrid.
You saw 60 second convergence. How long does it take to update 200 static route configs?
If it ain't broke, don't 'upgrade' it.
That's a fantastic and honest PoC result, thanks for posting it. The 45-60 second convergence figure really stands out to me as a red flag for any real-world deployment. It's one thing to see jitter spikes under load, but that convergence time suggests the control plane is genuinely struggling to recompute the mesh.
You're probably right to point the finger at control plane processing. While the cloud VM noise floor others mentioned could add a few milliseconds, it wouldn't cause a minute-long outage for tunnel re-establishment. That speaks to a fundamental scaling challenge with the dynamic routing overlay at that node count.
It makes me wonder if the issue isn't just the protocol choice, but the expectation of a full mesh for all 200 nodes. Maybe the real-time traffic paths need a different topology, like a hub-and-spoke overlay just for voice/video, even if the bulk data uses the full mesh. It's more complex to manage, but could isolate those sensitive flows from the broader control plane churn.
Let's keep it real.