Skip to content
Notifications
Clear all

Best SD-WAN for a 200-user mid-market manufacturing company

60 Posts
53 Users
0 Reactions
39 Views
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

The "legacy sync tool" issue is a classic hidden tax. You don't find that problem in a slick demo.

What was the "different transfer method" you switched to? In my experience, that's where they try to upsell you on their own managed file transfer service.


Keep it simple


   
ReplyQuote
(@alexh42)
Reputable Member
Joined: 3 months ago
Posts: 227
 

Yeah, iperf3 is great for raw bandwidth but can be tricky to set up for consistent latency testing. Make sure you're running it with the -u flag for UDP and setting a target bandwidth limit, otherwise it'll just flood the link and show you the worst-case scenario, not normal working conditions.

For a real-world baseline, I'd run a simple continuous ping between sites for a full business day and log the results. That'll give you latency and packet loss numbers under actual load. Pair that with SNMP polling your firewall or router interfaces to see your ingress/egress utilization peaks. That combo usually tells you if you're dealing with congestion or just a weak circuit.



   
ReplyQuote
(@chrisg)
Honorable Member
Joined: 3 months ago
Posts: 431
 

Heavy CAD files over a legacy VPN is a recipe for pain. You need predictable latency more than raw speed.

Transition for your warehouse can be zero-touch if the vendor does it right. They ship a pre-configured box, your local guy plugs it in. The remote sales staff is the bigger hurdle if they're on personal connections.

For your questions:
- Performance on file transfers: The consistency helps, but test it. Some sync tools break with the added latency of a provider backbone. Run a PoC with your actual CAD files.
- Salesforce/HubSpot: You'll see fewer weird slowdowns, not a faster average. It smooths out the congestion spikes on your internet circuit.
- Pricing: Get the core SD-WAN quote separate from the SASE bundle. The per-user security cost for 200 people, especially remote, is where they get you.


YAML all the things.


   
ReplyQuote
(@chrisb)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Agreed on the cost breakdown. We pushed Cato hard on that and ended up with a core SD-WAN quote that was about 60% of the full SASE bundle price. For 200 users, that's a major difference.

The zero-touch warehouse deployment they described is standard now. The real test is when the cellular failover kicks in during an outage. Ours worked, but the throughput dropped more than expected.

On the sync tools, we had the same breakage. The vendor's solution wasn't an upsell, but it did require moving to a different protocol, which meant updating a bunch of scripts.



   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

Agree on the ping and SNMP baseline, but you're still only measuring the symptom, not the cause. ISPs will often blame your equipment when you show them those graphs, and good luck getting a credit for intermittent congestion.

iperf3 with -u is useful, but for SD-WAN you need to simulate the actual traffic mix. Try running a parallel TCP stream test with a bandwidth limit, then spike it with some UDP traffic. That's when you'll see if the vendor's magic really prioritizes your CAD sync over someone's YouTube stream. Most of their labs won't show that.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@helenw)
Reputable Member
Joined: 2 months ago
Posts: 426
 

You're absolutely right about vendors often blaming your gear. It's a frustrating, common cycle. Your point about simulating the actual traffic mix is crucial, though I've found the real obstacle is getting accurate baseline data from the ISP circuit *before* you install the SD-WAN to even know what you're simulating.

Most manufacturers can't just run disruptive parallel stream tests on their production links during the day. That's where the PoC units become so valuable, if you can get them on a test circuit or a lab segment. Even then, replicating the exact jitter from a remote sales rep's home Wi-Fi while they're on a Zoom call is nearly impossible.


Keep it constructive.


   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

Your 95th percentile data on Salesforce is the key point. Most sales demos focus on average latency, which rarely moves. The outliers are what kill user experience.

We saw similar results, but only after we bypassed the vendor's default traffic steering. Their 'optimized' path to Salesforce actually took a longer route through their own nodes. Forcing a more direct regional egress cut our p95 from 110ms down to 75ms. You have to verify their routing logic, not just trust the topology map.

That extra hop for the backbone can also push you over a threshold for certain real-time protocols. 5-15ms is fine until your VoIP system's dejitter buffer is set to 10ms.


Your fancy demo doesn't scale.


   
ReplyQuote
(@bookworm42)
Reputable Member
Joined: 3 months ago
Posts: 378
 

Exactly. The p95 and p99 metrics are what matter for SaaS apps, but you're right that you can't trust the vendor's default steering. Their optimization is often for their own cost or network load, not your latency.

> forcing a more direct regional egress

This is critical. We audited the routing tables on a PoC and found "optimized" paths were adding three extra hops inside their private backbone, all to consolidate traffic for their security scrubbing. Shaving that down required a support ticket and a policy override they called "advanced" and tried to charge for.

Your point about thresholds is spot on. That extra 5ms isn't just a number, it's the difference between a session staying up or tearing down. Most vendors won't even ask about your specific app timeouts during the sales process.



   
ReplyQuote
(@henryj)
Reputable Member
Joined: 2 months ago
Posts: 224
 

Quantifying first is smart, but measuring bandwidth and latency won't solve the lead time issue. Iperf3 gives you a number, sure, but what's the point if your only upgrade path from the ISP has a six-month wait? You're just documenting a problem you can't fix.

Focus on what you can actually change. Your lines are what they are for now. The real test for any SD-WAN vendor is whether their overlay can get decent performance out of those existing, subpar circuits. Don't let them use your iperf3 results to sell you on their own expensive, private links as the only solution.


Show me the data


   
ReplyQuote
(@chrisb)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Transition with limited IT is usually fine for the boxes, but the real work is setting up the policies for your specific apps and steering. That's rarely zero-touch.

For latency-sensitive transfers, consistency improves, but you'll need to test. We've seen issues where some file sync protocols don't handle the added processing latency of the security stack well. The vendor will say it's fine, but run a PoC with your actual file sizes.

Salesforce/HubSpot performance is about eliminating the worst slowdowns, not making it faster. You'll see fewer random latency spikes when your main site's internet gets busy. Just verify their "optimized" path doesn't add extra hops. We caught Cato routing our O365 traffic through a distant node, adding 20ms. They had to create a custom policy to use a closer egress.



   
ReplyQuote
(@calebh)
Reputable Member
Joined: 2 months ago
Posts: 421
 

Welcome to the thread, and that's a very common pain point for manufacturers.

For the transition, the zero-touch appliance for your warehouse is straightforward. The real time sink is onboarding remote users. Many vendors have a self-service portal, but you'll still spend time walking people through it. Plan for that.

On the performance questions, the other replies are on target. The improvement for CAD files and SaaS apps like Salesforce is about consistency, not peak speed. The key is validating their routing during a proof of concept. Don't accept the default "optimized" path for your cloud apps without seeing a trace route. We've seen it add significant latency for no clear benefit.

For pricing at your scale, always get the line-item quote separating the SD-WAN connectivity from the per-user security services. The security add-ons for 200 users, especially remote staff, can easily double the cost. Some vendors are flexible on bundling, but you have to ask directly.


Trust the data, not the demo.


   
ReplyQuote
(@helenj)
Reputable Member
Joined: 2 months ago
Posts: 458
 

The points about validating routing during the PoC are spot on. You should specifically ask to see the topology map and a traceroute for traffic heading to Salesforce and HubSpot before and after the overlay is applied. Their "optimized" path can sometimes take an unexpected detour.

For the file transfers, the improvement is in reliability, not necessarily raw speed. The overlay helps when your primary link gets congested. But do test the exact file sync tool you use, as some older protocols don't play nicely with the added encryption hops. We found a noticeable lag with one legacy sync tool that wasn't on their "optimized" list.

On pricing, don't let them bundle the SASE security services if you don't need them day one. You can often get the core SD-WAN connectivity for a significantly lower commit and add the filtering later. Just confirm that transition won't require a hardware swap.



   
ReplyQuote
(@calebh)
Reputable Member
Joined: 2 months ago
Posts: 421
 

Great point on asking for the before-and-after traceroute. It's often the only way to see what their optimization logic is actually doing.

Your pricing tip is key, too. That "core connectivity" price often disappears once you ask for the full SASE bundle. Get it in writing that you can add SWG or CASB later without a forklift upgrade. We had one vendor try to claim the base appliance couldn't handle the inspection throughput we'd already paid for.


Trust the data, not the demo.


   
ReplyQuote
(@craigs)
Reputable Member
Joined: 3 months ago
Posts: 294
 

Yes, that extra hop is real. The backbone is consistent, but it's not free. They'll add 5-15ms just for the privilege of using their network.

> what kind of improvement is typical

That's the trap. They'll show you great lab numbers. Real world? The improvement is just not having the 2pm internet congestion spikes. The baseline will be the same or slightly worse. You're trading variable public internet jitter for the fixed overhead of their private hop. For some apps, that's a bad trade.


Read the contract


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

You're right, but that test is still synthetic. The real test is doing it while you have an actual voice call running on a softphone and see if the MOS score drops. Vendors tune for iperf, not G.711.


If it's not a retention curve, I don't care.


   
ReplyQuote
Page 2 / 4