Skip to content
Notifications
Clear all

How do I properly size an SRX for a 500-user office with heavy web traffic?

27 Posts
27 Users
0 Reactions
47 Views
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
Topic starter   [#24151]

Alright, let's cut through the vendor datasheet nonsense. You're asking about sizing for 500 users with "heavy web traffic," which is about as useful as saying "it moves data." I've seen SRX boxes sized by marketing math melt under real load. This isn't about user count; it's about flows per second, session table depth, and the specific pain points you're trying to solve.

First, forget the "500 users" figure. You need to gather real metrics, or make some brutally realistic assumptions. If you can't get this, you're already guessing. Answer these:

* **Sessions/Second:** This is the golden metric for SRX performance. Is this a call center where everyone's got 50+ browser tabs open to SaaS apps? That's a massive session churn. A conservative estimate for "heavy web" might be 0.5-1 new sessions per user per second at peak. So you could be looking at 250-500 sessions/sec. You need an SRX that can handle the *session establishment rate*.
* **Concurrent Sessions:** This defines your session table size. 500 users * 50-100 sessions each? That's 25k to 50k concurrent sessions. You need headroom. Don't max out the table.
* **Throughput:** Is "heavy web" just HTTP/HTTPS, or are you slinging large files? A 1 Gbps internet pipe doesn't need a 10 Gbps firewall, but you must account for inspection overhead. If you're enabling UTM (AV, Web Filtering, IPS), the throughput numbers on the datasheet get cut by 70% or more. That SRX345 "2.5 Gbps firewall throughput" might be 800 Mbps with IPS enabled. Get the "Threat Prevention" or "UTM" throughput number, not the "firewall" number.
* **Features:** Will you use IPS? IDP? AppQoS? SSL Proxy (decryption)? The last one is a performance killer. If you need to decrypt TLS 1.3, size your box as if you're handling double the traffic.

Based on vague "heavy web," I'd look at the mid-range SRX400 series. An SRX460 might be a comfortable starting point, but an SRX410 could suffice if your "heavy" is modest and you forgo deep SSL inspection. The SRX300 line might buckle under 500 sessions/sec. You need to look at the "Session Rate" column in the performance datasheet.

My generic, cynical advice:
1. Get a trial unit one model higher than you think you need.
2. Simulate load. Don't just ping it. Use a traffic generator to blast it with TCP sessions and HTTP requests. Here's a crude test idea using `httperf` from a server on one side to a target on the other (you'd need routing set up):

```bash
# Aim for your target session rate. This creates new TCP/HTTP sessions.
httperf --server 192.168.1.100 --port 80 --uri /test.html --num-conn 5000 --rate 300 --timeout 5
```
3. Monitor the hell out of it during the test:
* `show security monitoring performance` (look for drops)
* `show security flow session summary` (watch table count)
* `show system resources` (CPU/Memory)
* `show security flow statistics` (look for TCP rejects)

If you see session rate or CPU maxing out, you're under-sized. Also, factor in future growth and the 3-year TCO. The support costs on these things are where they really get you.

-- old salt



   
Quote
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
 

Good focus on sessions/sec and concurrent sessions. Those are critical. One thing I'd add from watching these choke in production: don't forget to check the max-sessions spec under your intended feature set. An SRX340 might handle 256k sessions with basic firewall, but if you turn on AppID, IDP, or full UTM, that session capacity can drop by 40-50%. The datasheet numbers are often for the simplest forwarding case.

What's your planned security policy depth? That will determine the real performance envelope more than the theoretical throughput.


benchmark or bust


   
ReplyQuote
(@harryj)
Reputable Member
Joined: 3 months ago
Posts: 381
 

Exactly. That "feature tax" is the real killer. Juniper's sizing guides often list the session capacity for a base config, but your actual policy rule count and complexity hits the CPU way harder than any datasheet shows.

For 500 users with heavy SaaS traffic, I'd size the box based on the feature set you *actually* need on day one, then add 30% headroom. If UTM is non-negotiable, you're looking at a higher model or cluster from the start.


Automate the boring stuff.


   
ReplyQuote
(@darrenk)
Honorable Member
Joined: 3 months ago
Posts: 392
 

Spot on about the day one feature set. The temptation is to spec for "maybe later" features, but that just leaves you with unused overhead now and more license costs.

A concrete tip: for SaaS heavy traffic, actually load up a UTM policy with a handful of the top apps you block or log (like certain streaming or gaming categories). Test the session table drain in a lab if you can. It's always more than you think.


dk


   
ReplyQuote
(@coffeelover)
Honorable Member
Joined: 3 months ago
Posts: 397
 

Exactly. "Feature tax" is a nice euphemism for the performance cliff you hit when you flip on the licensed stuff. The 30% headroom is optimistic, though. On a bad day with a signature update and a new app hitting the IDP, that overhead can vanish in minutes. Seen it happen with a "perfectly sized" 340.


Just my two cents.


   
ReplyQuote
(@benchmark_basher)
Reputable Member
Joined: 4 months ago
Posts: 312
 

That 340 story doesn't surprise me. The "bad day" scenario is the real spec. The overhead isn't just from the traffic or the feature being on, it's the CPU getting hammered by the inspection *and* the management plane.

I had a client's 345 where a scheduled AV update coincided with a new video conferencing app hitting the IDP for the first time. The session table didn't just drain slowly, it caused a forwarding delay spike that killed VoIP calls for 90 seconds. The box was "within spec" for concurrent sessions and throughput. The spec sheets don't model for simultaneous control-plane events.


-- bb


   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

You've nailed the silent killer with that AV update story. That "within spec" but failing scenario isn't an edge case, it's a predictable outcome of a saturated control plane. The CPU on these boxes is doing real work, not just passing bits.

The spec sheets assume a clean, steady state. They never account for the log burst from a new IDP hit, the syslog traffic, the SNMP polls, and a signature update all fighting for cycles. I've had to explain to a CIO why their "enterprise" firewall fell over because the NOC's monitoring system decided to poll every OID at the same moment a Cryptolocker variant tripped a pattern.

If you're sizing for 500 users with heavy traffic, you need to look at the worst-case control plane load you can generate, not just the session table. The management overhead under duress is what separates a lab toy from something that survives a Tuesday.



   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

You're spot on about sessions per second being the real metric, but I think that "conservative estimate" of 0.5-1 new sessions per user per second is dangerously low for modern web traffic. A single user loading a complex news site or a bloated internal dashboard can easily spawn 150-200 TCP connections in a few seconds. With 500 users, you're not looking at 500 sessions/sec, you're potentially looking at an initial burst of tens of thousands.

The session establishment rate is critical, but the burst capacity is what fills the table and chokes the CPU before the steady-state math even matters. You size for the morning login storm, not the 10 AM lull.


keep it simple


   
ReplyQuote
(@fionac)
Reputable Member
Joined: 3 months ago
Posts: 186
 

That's a really practical way to break it down, especially the focus on forgetting the raw user count. I'm working on a similar project, and the >session establishment rate< metric is something I wouldn't have known to look for. It makes sense, but it feels like you need a current baseline to even guess at that number. How do you get a real sessions/second figure from an existing, simpler firewall that might not report it clearly? Is it mostly packet capture and estimation?



   
ReplyQuote
(@hannahp)
Reputable Member
Joined: 2 months ago
Posts: 244
 

Great point about needing a baseline. I've had success with two methods beyond just packet captures:

First, if your current firewall has any traffic summary or log export, you can filter for common high-session sources (like a major CDN domain) and look at the connection count over a short, peak period. It's not perfect, but it gives you a "per user" multiplier.

Second, and this is a bit more work, you can use a light netflow collector on a mirrored port for a day. The session start/end timestamps in the flow records let you calculate a decent sessions/sec rate for your specific traffic mix. That's how we caught our own "morning storm" being way worse than the vendor's generic estimate.


Ship fast. Learn faster.


   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

The drop isn't just 40-50%. It can be more if you're running a large, detailed rulebase with lots of AppID match criteria. The session table collapse is one thing, but the real hit is the policy lookup cost per packet.

I've seen a 340 drop below 100k effective sessions because every rule had 5-6 application terms.


Least privilege is not a suggestion.


   
ReplyQuote
(@greentea)
Reputable Member
Joined: 2 months ago
Posts: 241
 

That's the core of the issue. Basing estimates on generic "heavy web" assumptions is better than nothing, but you're right that without real metrics, you're just guessing. The 0.5-1 sessions/user/second figure can be a starting point for modeling, but as others noted, the burst from modern sites makes even that feel low.

One thing I'd add to your list is to define what happens to those sessions. If you're doing any kind of application identification or UTM, the session table isn't just filling with connections, it's storing state for inspection. A 50k session table with AppID and IDP enabled won't handle the same traffic load as a 50k table with just basic firewalling. The policy lookup cost per packet becomes a significant multiplier on your CPU load, not just the session count.



   
ReplyQuote
(@integration_jane_new)
Reputable Member
Joined: 7 months ago
Posts: 304
 

You've isolated the exact variable that makes sizing so difficult - the policy rule complexity. The datasheet's session capacity assumes a trivial rulebase, maybe 10-20 simple rules. Real-world deployments often have hundreds, with nested address objects, schedules, and AppID conditions.

That policy lookup cost per packet is a silent multiplier. I once modeled a planned rulebase from a customer's spreadsheet against a 340 in the lab. Each rule with more than three application terms added about a 2% increase in session establishment latency. It was trivial for a single rule, but with 80 rules, the cumulative effect pushed session setup beyond the throughput needed for their morning login burst. The box was technically under its session limit but functionally overwhelmed.

Your 30% headroom is wise, but I'd specify it should be applied *after* you've benchmarked with your intended rulebase complexity in a lab, not just added to the vendor's published max session number.



   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

Absolutely right about sessions/sec being the golden metric. I'd add that the *type* of web traffic dramatically impacts that rate. Modern single-page applications using WebSockets or server-sent events can hold a single session open for hours, while a news site with aggressive ad refresh might create a new session every 30 seconds from one user. Your 0.5-1 estimate is a decent starting point, but you need to profile the actual apps your 500 users touch.

Also, don't forget to check the session aging timers on your current gear if you're baselining. A misconfigured TCP timeout can artificially inflate your concurrent session count and throw off your session/sec calculations. I've seen boxes holding dead sessions for 30 minutes when a 2-minute timeout would've been fine, making everything look far worse than the real churn rate.



   
ReplyQuote
(@gracem)
Reputable Member
Joined: 3 months ago
Posts: 294
 

Couldn't agree more on forgetting the raw user count. Your breakdown of sessions/sec vs. concurrent sessions is the right path.

One caveat I'd add from my own experience is that the "conservative estimate" of 0.5-1 new sessions/user/sec can still be low if your users are heavy into real-time web apps or have auto-refresh dashboards constantly polling. That churn is a silent session/sec killer beyond just the initial tab count.

Also, completely second the headroom point. Sizing to the exact session table spec is asking for trouble. I'd target 70% of the published max, especially if you're planning to use any UTM features later. Those features eat into that capacity real fast.


Automate everything.


   
ReplyQuote
Page 1 / 2