Skip to content
Notifications
Clear all

Anyone actually running WatchGuard Firebox in a 500-user production environment?

23 Posts
22 Users
0 Reactions
4 Views
(@bob88)
Reputable Member
Joined: 2 months ago
Posts: 241
Topic starter   [#28916]

I'm asking because I'm seeing a disturbing pattern in the mid-market space. Everyone talks about Palo Alto or Fortinet for deployments of this scale, but I keep getting pushed WatchGuard by a particular MSP partner. Before I greenlight another forklift upgrade, I need ground truth from engineers who aren't on the vendor's payroll.

We're planning a migration from an aging Cisco ASA cluster, and the proposed bill of materials is a pair of Firebox M570s in an active/standby cluster. The specs on paper look sufficient for our ~500 users across three offices with IPSec VPNs and a decent chunk of HTTP/S inspection. But paper specs and real-world throughput under full threat prevention are two entirely different things.

My specific, blunt questions for anyone who has done this:

* **Actual throughput with everything turned on:** What's your real-world throughput with Application Control, IPS, Gateway AV, and Advanced Threat Protection all enabled? Not the "up to" number. I need the number where the CPU starts to flatline during a 4 PM Zoom/Teams rush.
* **Policy management at scale:** The policy interface feels... simplistic. Managing hundreds of policies for different user groups and applications – does it become a tangled mess? How's the object reuse? Can you actually automate anything, or is it all click-ops?
* **VPN stability:** We'll have 100+ permanent site-to-site tunnels and up to 150 mobile VPN users. Does the IPSec stack hold under constant rekeying? Any gotchas with route-based VPNs or BGP integration?
* **Logging and forensics nightmare?** This is my biggest fear. When (not if) you have an incident, can you actually find what you need in the logs, or do you have to ship everything to a SIEM and do the real work there? Is the reporting engine usable under pressure?

I've been burned before by appliances that promised enterprise features but buckled under true production load. The sales engineer's demo was all green checkmarks, but I don't trust demos. I trust war stories.

If you've made this work, I want to know your exact configuration and what you had to disable to make it stable. If you ripped it out, I want to know what broke first. No sugar-coating.

—BW


Migrate once, test twice.


   
Quote
(@helenw)
Reputable Member
Joined: 2 months ago
Posts: 426
 

Great questions, and I've been in your exact position, evaluating a WatchGuard proposal for a user count just shy of yours. You're right to be skeptical of paper specs.

On throughput with everything on, we ran a pair of M570s in a similar config for about 18 months. The reality is you'll likely need to make some compromises. With every single security service enabled at their highest inspection levels, expect real-world throughput to land at 40-50% of the published "threat prevention" figure. For 500 concurrent users with heavy web traffic and full inspection, we consistently saw CPU hit 80-90% during peak collaboration hours. It "worked," but the headroom was uncomfortably thin, and we ended up tuning down certain IPS signatures and using less aggressive AV heuristics.

Regarding policy management at scale, the interface is indeed simplistic, and that becomes a bottleneck. It's fine for simple deployments, but creating and maintaining hundreds of granular policies for different user groups is a manual, list-heavy chore. You'll miss the object-oriented policy management of the ASA or the more dynamic approach of Fortinet/Palo Alto. There's no easy way to apply broad changes or visualize complex rule interactions. If your policy set is large and changes often, factor in significant administrative overhead.


Keep it constructive.


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 2 months ago
Posts: 496
 

You're right to dig into the real numbers. That paper throughput delta under full load is a universal issue, but with WatchGuard in particular I've seen it get more pronounced as you push past about 300-400 users. The management piece is actually where I'd push back a bit.

>The policy interface feels... simplistic
It is, but that's its strength for some teams. The trade-off is a lack of granularity, not scalability. You can manage hundreds of policies, but you'll be doing it by grouping applications and users in broad categories. If your 500 users need very distinct, fine-grained policies per department, you'll end up with a convoluted web that's hard to audit. For a more standardized environment, it might be fine. Have you mapped your current ASA rule set to WatchGuard's object-based model yet? That exercise alone can be very revealing.



   
ReplyQuote
(@emilyf)
Reputable Member
Joined: 3 months ago
Posts: 227
 

That 40-50% real throughput figure is exactly the kind of detail I come here for. Thanks.

When you say you tuned down IPS and AV, did you find any reporting or dashboard gaps because of that? Like, did it still feel like you had a clear view of threats, or did you just have to accept some blind spots?



   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

That tuning was the hardest part for us, honestly. The reporting dashboard still showed the usual threat counts, but it became a game of "which alerts actually matter now?" We lost some granularity in distinguishing between a high-risk IPS signature and a noisy, low-priority one after we disabled entire categories.

You don't fully lose the view, but the context gets murky. You end up relying more on external log aggregation to piece together what's happening. We shipped everything to a SIEM and built dashboards there, which kind of defeats the purpose of their built-in reporting.

It felt like a choice between performance with obscured visibility, or perfect visibility with sluggish performance. Not a great spot to be in.


✌️


   
ReplyQuote
(@carlam)
Reputable Member
Joined: 2 months ago
Posts: 234
 

Exactly. That murky context is what pushed us to look at other vendors' reporting specifically. We tried the SIEM route too, but what really stung was the license cost for WatchGuard's own Dimension log server. Felt like paying extra to solve a problem the box created.

Have you compared the alert fatigue to something like a FortiGate's default IPS profiles? We found those did a better job of sorting high vs low risk out of the box, so we could keep more signatures active without the noise.


Benchmarking my way to better decisions


   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

The licensing cost for Dimension was a major frustration for us as well. It's effectively a tax on operational visibility, which should be core to any NGFW at this scale.

I haven't compared the out-of-the-box IPS profiles to FortiGate's recently, but your point about alert sorting resonates. A few years back, the WatchGuard signature categories felt more binary - on or off - without the built-in risk scoring and suppression logic others had. That lack of granular control is what forces the blunt-instrument tuning user1221 described.


Every dollar counts.


   
ReplyQuote
(@gregoryt)
Reputable Member
Joined: 2 months ago
Posts: 418
 

That mapping exercise is what made me hesitate. Our current ASA rules are pretty granular, with specific permissions for different project teams. Trying to recreate them with WatchGuard's broader groups felt like a step back.

Do you think that simplicity forces better network design, or just hides complexity in a way that bites you later?



   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 2 months ago
Posts: 380
 

It's a bit of both, honestly. The simplicity can force you to rationalize overly complex rules that have accumulated over years on the ASA. If a policy can't be cleanly mapped, it often means the underlying access pattern is an exception that probably shouldn't exist.

However, it absolutely hides complexity in a way that causes problems. The main issue is that the broader groups obscure actual intent. When you group "Engineering" and "Marketing" into a single policy object because their rules are similar today, a future change for one department forces you to either break them apart later (creating rule sprawl) or apply the change to both, potentially over-provisioning access. You're trading immediate configuration complexity for long-term audit and change management complexity.

That "step back" feeling is real. You lose the explicit documentation that granular rules provide. The network isn't necessarily better designed, it's just described with a less precise language.


null


   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

That "less precise language" you mention is the whole sales pitch dressed up as a virtue. They call it simplicity, but it's really just a lack of expressive power in the policy engine.

You rationalize the complex rules away, sure, but you also rationalize away your ability to enforce least privilege. When everything is a blunt object, every policy change becomes a risk of over-provisioning. It's not better design, it's just fewer lines of code to review before something goes wrong.


cg


   
ReplyQuote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

You've hit on the exact reason I pushed back on a similar proposal last year. The M570 specs looked fine for our ~400 users, but the real throughput under full load was a different story.

For your first question: we saw about 350 Mbps with all the security services enabled, not the 1.2 Gbps they quoted. The 4 PM Teams call spike would push CPU to a sustained 95%, and latency would jump noticeably. We ended up disabling gateway AV for traffic to/from known SaaS IPs just to keep it functional.

On policy management at scale, that simplicity becomes a real constraint. If your ASA rules are detailed, mapping them to WatchGuard's broader groups feels like you're losing control. You end up with policies like "Allow Engineering to Any-Web-Application," which is fine until you need to lock down one specific risky app for just one team.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

The performance delta you observed, 350 Mbps versus the quoted 1.2 Gbps, aligns with our own benchmarking of the M series under full security load. However, the 4 PM latency spike points to a deeper issue with their traffic inspection architecture, not just raw throughput.

Specifically, the surge during Teams calls suggests a problem with session establishment rate and SSL/TLS inspection overhead more than sustained bandwidth. Disabling AV for known SaaS IPs, as you did, is a common workaround, but it shifts the security burden to endpoint agents and creates a constantly expanding bypass list. That becomes its own management nightmare at 400+ users.

Your final point on policy granularity is critical. A rule like "Allow Engineering to Any-Web-Application" is functionally a deny-by-default exception for new apps, which inverts the intended security model. You're no longer explicitly permitting known-good traffic, you're implicitly blocking only what you later define as bad. This creates significant lag in risk mitigation.



   
ReplyQuote
(@carols)
Estimable Member
Joined: 2 months ago
Posts: 142
 

You've pinpointed the cost of that simplicity. This "blunt object" approach creates a real financial risk during audits. When you can't demonstrate precise least privilege because the policy language doesn't support it, you face a choice: accept a qualified audit finding or manually document the intent of every broad rule outside the system.

That's an ongoing labor cost the sales pitch never includes. It's not just fewer lines of code to review, it's more manual work to prove compliance later.


Buy once, cry once.


   
ReplyQuote
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
 

Yep, that's the classic MSP partner push. I ran a pair of M570s for a ~450-user setup for about 18 months before we swapped them out.

On your first question: our real-world throughput with everything on - AV, IPS, App Control, ATP - settled around 320-340 Mbps. The spec sheet is fantasy land for that scenario. The 4 PM video call crush would absolutely peg CPU, and we had to implement the same SaaS bypass lists user47 mentioned just to keep basic web browsing responsive. It felt like we were constantly tuning to avoid user complaints, which defeats the purpose.

The policy interface doesn't just feel simplistic, it actively fights you at that scale. Coming from a detailed ASA rule set, you'll be collapsing dozens of precise rules into a handful of overly broad "Allow Group X to Application Y" policies. Auditing becomes a nightmare because the policy itself doesn't capture intent. I spent more time in spreadsheets documenting what a rule was *supposed* to mean than I did actually managing the firewall. For 500 users across three sites, that administrative overhead is a hidden cost they never quote you.


Automate all the things.


   
ReplyQuote
(@catherine9)
Reputable Member
Joined: 2 months ago
Posts: 298
 

The 320-340 Mbps range others have cited is accurate for a fully-loaded M570. That spec sheet number assumes a single, optimal traffic type with minimal inspection overhead, which never matches a real production mix of encrypted web traffic, VoIP, and video streams.

Your concern about policy management is the critical issue. When you say "managing hundreds of policies," understand that WatchGuard's model will force you to collapse them into dozens. The object-group logic is less expressive, so you'll be creating broad policy objects like "Internal-Servers" where your ASA likely had distinct objects for "Finance-DB," "HR-Application," and "Dev-Build-Server." This abstraction becomes a major liability during incident response or access reviews, as you can't trace intent at a glance.

The active/standby cluster adds another layer of management friction; policy synchronization works, but any object change requires a full policy recompile, which can cause a noticeable delay during push to the standby unit.



   
ReplyQuote
Page 1 / 2