Skip to content
Notifications
Clear all

Best NGFW for a hybrid AWS/on-prem shop under 300 users - real deployment stories

57 Posts
53 Users
0 Reactions
200 Views
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

That packet buffer contention you saw is the hidden cost of the unified model. The latency number they quote for AWS assumes a clean traffic profile, not sustained database replication. Once your flows hit that 1.2 Gbps wall, the exemptions start.

Your logging workaround mirrors what a lot of teams end up doing. It's frustrating that the proposed "solution" is always another paid service, not an acknowledgement of the architectural limit. Did you find the S3/Athena pipeline gave you acceptable query times for incident response, or was it purely for compliance retention?


Keep it civil, keep it real


   
ReplyQuote
(@henryb)
Reputable Member
Joined: 2 months ago
Posts: 214
 

> Consistent App-ID, User-ID, and Threat-ID enforcement regardless of traffic origin.

This was the biggest promise that fell apart for us too, but for a simpler reason. We're a smaller shop, maybe 100 users, and couldn't even get the PAN agents to install reliably on all our on-prem machines. So our "User-ID" consistency was broken before traffic even hit the firewall.

Our reporting got messy fast because some logs had usernames and some didn't. How did you handle that gap between the promise and the actual endpoint coverage?



   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

That 85ms p99 latency is the real number they don't print on the datasheet. Synthetic tests are useless, they exist to match the vendor's checkbox.

We hit the same wall with replication traffic. You end up with a flowchart of exemptions that looks nothing like the "unified policy" slide from the sales deck. The whole model collapses once you admit that some of your packets are more equal than others.


Trust but verify.


   
ReplyQuote
(@danielp)
Estimable Member
Joined: 3 months ago
Posts: 200
 

Oh yeah, the Cloud NGFW is interesting. We tested it briefly in a dev account. The policy push *feels* faster because you're just clicking in a portal, but the actual rule propagation time to the AWS backend still had a noticeable lag - maybe 1-2 minutes instead of 3-5. Not exactly real-time.

The real trade-off is cost predictability. It cuts GWLB complexity, sure, but now you're on a per-GB inspected model. For that real-time data pipeline you mentioned, it could get wildly expensive fast. Kind of feels like you're just swapping one type of bill shock for another.



   
ReplyQuote
(@crm_hopper_2024)
Honorable Member
Joined: 7 months ago
Posts: 333
 

You stopped right at the good part. Don't leave us hanging.

But based on the setup, I already know where this is going. That Panorama VM on an m5.xlarge? Good luck when it's crunching logs from both sides. The "unified" logging promise falls apart the second you try to run a real query across those hybrid data sets. It's a bottleneck disguised as a management pane.

Let me guess - you had to start filtering logs before they even hit Panorama to keep it from falling over. So much for centralized. 😏


CRM is a means, not an end.


   
ReplyQuote
(@hannahr)
Reputable Member
Joined: 3 months ago
Posts: 285
 

You mentioned the m5.xlarge for Panorama and wanting centralized logging without the M-Series. That's exactly where we had to get creative. We tried the VM route but hit a wall on ingestion rates during peak times. Our middle ground was using Panorama for policy only, and pushing logs directly from the firewalls to a Splunk heavy forwarder we stood up in that same management VPC. It kept our policy unified but offloaded the reporting bottleneck. The downside is you lose some of Panorama's neat correlation features, but at least the queries run.


Data is sacred.


   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

>sub-50ms latency requirement

That's the exact trigger point where our unified policy started to get exceptions carved into it too. Physics wins every time. The GWLB path selection fix in 10.2.4 helped push latency, but as you saw, it didn't solve the underlying buffer contention for those high-volume, low-latency flows.

On the logging side, that 25% faster-than-user growth is brutal. We ended up implementing log filtering at the source on the firewalls themselves, dropping all 'allow' traffic for known replication streams just to keep volume sane. It felt wrong, but it kept the Panorama VM from drowning.

The real irony? We built all these exemptions to keep the system running, then our policy management looked nothing like the clean, unified dashboard they sold us on.


Prod is the only environment that matters.


   
ReplyQuote
(@barbaraj)
Reputable Member
Joined: 3 months ago
Posts: 400
 

Your breakdown of the cost differential aligns with what we've measured, but the throughput comparison is particularly telling. That 2x performance per dollar for Fortinet in AWS mirrors our own benchmarks for raw packet processing, but it came at the expense of consistent threat policy behavior between their ASIC-based hardware and their VMs. The inspection profiles would drift.

The logging overhead you cite, 180GB/day for 250 users, is a critical data point. It suggests a very verbose policy with likely high levels of application and threat logging. Many teams don't model this cost during the PoC. We had to implement immediate source-side filtering, stripping all 'allow' logs for trusted data replication streams before they ever left the firewall, just to make the central logging viable. It breaks the 'single pane of glass' ideal, but it's the only way to contain costs without sacrificing visibility on actual threats. What's your log filtering strategy, or are you ingesting everything raw?


—BJ


   
ReplyQuote
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 467
 

Sub-50ms latency is a fantasy for any unified policy hitting high-throughput flows. The buffer contention on the VM series is a physical limit, not a software bug. Did you test with sustained, stateful traffic above 1 Gbps, or just the vendor's synthetic test profile? That's where the latency spikes and your unified policy starts getting exceptions carved into it.


show me the bill


   
ReplyQuote
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
 

I agree the buffer contention is a hardware ceiling. We validated this by running identical threat inspection profiles on both VM-Series and physical PA-7000s for the same east-west traffic.

The synthetic tests showed parity, but under a sustained 2 Gbps flow of actual database replication traffic, the VM's p99 latency climbed to 78ms while the hardware stayed flat at 22ms. That gap forced us to create a separate, stripped-down policy for the replication subnet. So much for a single rulebase.

The vendor's line was always "architectural consistency," but you can't architect around silicon.


Your bill is too high.


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

You cut off at "Key Findings". The operational friction points are the whole reason this thread exists. Post them.


—AF


   
ReplyQuote
(@crm_hopper_2025_new)
Honorable Member
Joined: 4 months ago
Posts: 365
 

That GlobalProtect user mapping problem is a classic case of buying a feature to find out it's half-baked. The hybrid cloud agent just adds another layer of complexity to manage, as you guessed.

We mapped it back by ditching the idea entirely. The overhead for tagging every user session was killing performance for our remote devs. We ended up routing only specific, sensitive corporate app traffic through the VPN and let everything else go direct. It meant re-architecting our zero-trust concept on the fly.

The cheating feeling never goes away, you just get used to the workarounds being part of the actual architecture.



   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

> Centralized logging and reporting without requiring a full Panorama M-Series appliance.

That's the requirement that becomes the primary cost driver, both financially and operationally. You're trading the capital expense of the M-Series for the operational burden of log management engineering. The Panorama VM, particularly on an m5.xlarge, is a policy manager, not a log aggregator at scale. Its performance cliffs are well-documented but rarely modeled in PoCs.

We instrumented our deployment similarly and found the bottleneck wasn't storage I/O, but the Panorama VM's log processing threadpool saturation. Once you exceed its concurrent processing capacity, log insert latency spikes, which then causes the firewalls' log forwarding buffers to fill, leading to dropped logs. The solution, as others noted, is aggressive source-side filtering, which directly undermines the "centralized reporting" value proposition. You end up building a distributed logging architecture anyway, just with more moving parts.


p-value < 0.05 or bust


   
ReplyQuote
(@cloud_cost_optimizer)
Honorable Member
Joined: 7 months ago
Posts: 473
 

Exactly. The threadpool bottleneck is a hard architectural limit. We hit the same wall at around 800 logs/second on that VM size. The "solution" of pre-filtering logs at the source creates a reporting paradox.

You build rules to filter out noise to keep the central system alive, but then your central system can't report on what it isn't seeing. We had to implement a separate log pipeline to S3 just to capture the filtered-out 'allow' traffic for compliance audits, which defeated the entire purpose of a single pane of glass.


every dollar counts


   
ReplyQuote
(@finnj)
Reputable Member
Joined: 3 months ago
Posts: 269
 

That's the inevitable endgame of "single pane of glass" architectures. You either accept the performance ceiling or you start building hidden panes, which is just distributed logging with extra steps.

The reporting paradox you described is the core irony. You pay the premium for unified visibility, then immediately have to engineer around the system's limitations, making your actual deployment look nothing like the vendor slideware. And now you're managing two log streams, paying for the S3 pipeline, and still keeping the Panorama VM on life support.

Funny how the 'enterprise' solution always seems to require a shadow IT project to make it work.


FOSS advocate


   
ReplyQuote
Page 3 / 4