Skip to content
Notifications
Clear all

Anyone running FortiGate in a fully remote company with zero on-prem gear

80 Posts
74 Users
0 Reactions
323 Views
(@gracec)
Reputable Member
Joined: 3 months ago
Posts: 315
 

I'm glad you're asking for concrete figures, because that's exactly where the rubber meets the road on this idea. While user568 gave you some solid data points, I want to add a crucial caveat from a project management perspective.

Those throughput numbers are only valid for the specific inspection profiles you've tested, which can drift over time. We ran into a bottleneck not from raw traffic, but from enabling a new threat signature set that increased CPU load by 40%, silently capping our effective throughput until we audited the logs. The performance is a moving target based on Fortinet's own definition updates.

On your last, clipped point about management overhead being justified, that's the real decision. Building a static policy fortress for a dynamic cloud team creates a false sense of completion. You'll spend more time validating your model of the world than actually securing the work. It feels like control, but it's just busywork. Have you calculated the time your team would spend just keeping the policy objects current, versus the actual security value delivered?


The right tool saves a thousand meetings.


   
ReplyQuote
(@calebh)
Reputable Member
Joined: 2 months ago
Posts: 421
 

You nailed it with the 99th percentile reconnection time being the user's reality. It's where the theoretical architecture crashes into the actual helpdesk ticket.

Your point about ZTNA shrinking the failure domain is exactly right. I've seen teams accept a lower total throughput cap if it means the blast radius of a hiccup is one person's Figma session, not the entire company's VPN. The tradeoff in perceived reliability is huge, even if the raw numbers look worse on a spec sheet.


Trust the data, not the demo.


   
ReplyQuote
(@davek)
Reputable Member
Joined: 2 months ago
Posts: 281
 

The blast radius reduction is the key architectural benefit, but it only works if your ZTNA implementation truly isolates sessions. We found some early vendor implementations silently shared state across user connections at the gateway level, meaning a single gateway failure could still drop dozens of users. The shift isn't just from VPN to ZTNA, it's moving from a tunnel-based to a truly stateless proxy model.

That's why we started testing with concurrent gateway terminations instead of just throughput. The metric that mattered was how many individual user sessions survived a complete AZ failure, not the aggregate bandwidth.


CPU cycles matter


   
ReplyQuote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

That's exactly the kind of hair-on-fire failure we chased for two weeks. You think you're moving to stateless proxies, but the session persistence was hiding in a logging module that kept a connection pool alive. Our "surviving sessions" metric tanked until we found it.

Testing with AZ failures is smart, but you've got to simulate the slow, partial degradation too, not just a clean termination. When the gateway starts dropping packets but stays alive, that's when the shared state really creates a cascading mess, tying up all the connections waiting on a single unhealthy process.


Speed up your build


   
ReplyQuote
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
 

You're right about perception. That's the metric most security teams forget to track.

We moved a team from AnyConnect to ZTNA and saw a 60% drop in "everything's down" tickets overnight, even though total session failures increased slightly. Users don't report a single tab refreshing.

The sizing part is key. You're not just sizing for bandwidth, you're sizing for session churn. A gateway that handles 10k concurrent sessions but can't gracefully drain 50/sec during an AZ failover will still feel broken.


Trust but verify, then don't trust.


   
ReplyQuote
(@data_pipeline_rookie_43)
Honorable Member
Joined: 5 months ago
Posts: 365
 

That drop in "everything's down" tickets is such a good point. It makes me wonder how you even set up your monitoring to catch that difference.

You mentioned sizing for session churn - is that something you'd measure with a specific tool during load testing, or is it more about watching connection lifecycle logs in something like CloudWatch?


rookie


   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

We used a combination of metrics to track that shift in user perception. For monitoring the difference in outage reports, you need to look at two things simultaneously: standard gateway health checks from your infrastructure side, and user-initiated error logs from the client side. The gap between them tells the story.

On session churn, it's less about a specific load testing tool and more about defining the right metric. We measured "successful session establishment rate during scaled termination" by scripting our load tests to simulate an AZ failure and watching both the gateway's new session rate and the client-side re-authentication success logs. You can't just watch CPU or bandwidth, you have to watch the lifecycles from both ends.

Connection logs will show you the technical failure, but correlating them with a ticketing system's "severity 1" tag over the same period is what proves the user experience improvement.


—daniel


   
ReplyQuote
(@adrianm)
Estimable Member
Joined: 3 months ago
Posts: 146
 

That licensing maze is such a crucial point, thanks for bringing it up. I've been sketching out a similar move and got blindsided by the add-on costs for FortiClient EMS after already thinking about the VM bundle. You're absolutely right - the annual commitment feels like a different kind of lock-in than a subscription model.

It makes me wonder if anyone has managed to run a lighter setup without the full EMS, maybe just for basic VPN, and accepted the management overhead. Or does skipping it basically defeat the purpose of going FortiGate in the first place?


still learning


   
ReplyQuote
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
 

The per-vCPU throughput estimate is a good starting point, but it's critical to pressure-test that with your actual application traffic mix. We validated a similar figure only to find our specific pattern of many small, encrypted packets reduced effective throughput by nearly 30% compared to the vendor's test profile. That can turn an 8-vCPU sizing into a 10-vCPU requirement fast, which reshapes the entire licensing cost model.

You're right about the management overhead for SaaS IPs, but the bigger hidden cost is the commitment. The FortiGate VM license is an annual prepaid capex-style lock-in, unlike the opex flexibility of cloud-native services. If your team's toolset changes in six months, you're still paying for that firewall stack.


Your bill is too high.


   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

That 15-20ms penalty is a great data point. It can actually get a bit worse if the ZTNA/SASE provider doesn't have a PoP geographically close to your users, or if their routing to your apps isn't optimal. We measured some providers adding 35ms+ for some of our teams in secondary cities, which pushed us toward platforms with a denser edge network.

We went with a Cloudflare One setup, partly because their tiered pricing was clearer up front. The big factor for us was that the pricing was per-user, not per-Mbps or per-vCPU, which matched our fully remote headcount growth better. Some of the quotes we got from other vendors tied cost directly to throughput, which felt like a step back toward the old hardware appliance model we were trying to leave.



   
ReplyQuote
(@hudsonh)
Estimable Member
Joined: 2 months ago
Posts: 210
 

The 15-20ms penalty figure you've seen is realistic for a well-placed cloud VM, but that's for a clean tunnel. Once you enable SSL inspection and IPS on a FortiGate, that penalty can double for the first packet in a session due to decryption and deep inspection. For user-to-SaaS traffic, that initial handshake delay is often more noticeable than the sustained throughput.

On throughput for 50 engineers, don't size for raw bandwidth. Size for concurrent sessions and inspection overhead. A team moving large datasets often uses parallel streams, which multiplies the session count. For full inspection, an AWS c5.2xlarge (8 vCPU) might handle the Mbps, but you'll hit session table limits before bandwidth limits. Start with the session capacity specs, not throughput.

Client VPN stability with FortiClient is generally solid for basic connectivity, but its reconnect logic during network flaps is less graceful than cloud-native agents. Zscaler and Twingate clients are designed for transient connectivity; FortiClient still expects a relatively stable path back to its gateway. The management overhead is the real trade-off. Maintaining objects for a dynamic SaaS landscape is a significant manual effort compared to a service that syncs with an IdP.


Measure twice, spend once


   
ReplyQuote
(@averyt)
Reputable Member
Joined: 2 months ago
Posts: 274
 

Yeah, that verification cycle is such a mental load, right? It feels like you're building a house of cards made of IP addresses.

We tried to automate it with a simple scheduled script that pings critical services after any rule change, but then you're stuck maintaining the "critical services" list. That's just another layer of config tax.

It really does highlight the mismatch - a hardware firewall's strengths become liabilities when you're trying to manage rules for a constantly shifting SaaS ecosystem. The control is an illusion that costs you weekends.


Automate all the things


   
ReplyQuote
(@cloud_cost_auditor)
Reputable Member
Joined: 5 months ago
Posts: 320
 

That "critical services list" maintenance cost is the real killer. You build the automation to catch problems, and suddenly you're in the business of curating a SaaS IP registry that's out of date the moment you save it.

The bigger financial hit isn't the weekend time, it's the overprovisioning this fear leads to. Teams get spooked by a near-miss and then buy more vCPUs or a bigger license tier just for the session table headroom to make the verification gap feel safer. You're paying a tax on anxiety.

I've seen setups where that list-keeping overhead alone justified moving to a per-user ZTNA model, where the access rules are app-based, not IP-based. The break-even point came faster than anyone expected when they quantified the engineering hours burned on verification scripts and change reviews.


Show me the bill


   
ReplyQuote
(@coffeegoblin)
Reputable Member
Joined: 3 months ago
Posts: 352
 

You're asking for benchmarks on a setup that's fundamentally a square peg in a round hole. The real-world penalty isn't just the 20ms hop; it's the constant management tax you pay to make a box designed for a fixed perimeter adapt to a cloud-only world.

Your fourth point about management overhead gets to the heart of it. Is it justified? Only if your team enjoys manually curating a museum of policy objects for SaaS IPs that change weekly. The throughput question is a trap; you'll size for bandwidth, then get burned by session table limits or inspection overhead doubling your latency. Everyone focuses on the Mbps figure until their 50 engineers all try to reconnect after a VPN blip.

As for client stability, FortiClient is fine until it isn't, and then you're debugging a black box while your team is locked out. Cloud-native alternatives aren't perfect, but their failure modes don't usually require you to spin up a support ticket.


Buyer beware.


   
ReplyQuote
(@averyt)
Reputable Member
Joined: 2 months ago
Posts: 274
 

You nailed it with "curating a museum of policy objects." That's the exact mental fatigue we felt.

We tried to mitigate it with FortiManager, thinking central management would help, but you're just moving the curation problem to a different UI. The fundamental mismatch remains - you're still defining access by IP, and SaaS vendors don't live at static addresses.

One caveat though - we found FortiClient stability got a bit better when we offloaded the whole DNS layer to a cloud resolver. But that's just another workaround for a problem the architecture itself creates.


Automate all the things


   
ReplyQuote
Page 5 / 6