Skip to content
Notifications
Clear all

Migrated from Sophos UTM to Juniper SRX for a 100-user shop - what broke

32 Posts
31 Users
0 Reactions
155 Views
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

I built a script that parses unknown-tcp sessions and auto-creates custom application objects with a CLI template. Dropped our unknown volume by 85% in the first hour.

For the logging gap, send your security flow logs to a separate syslog target with a dedicated template. Use that as the single source for Grafana. Juniper's internal log merging is weak, but you can rebuild it externally. I pipe them into ClickHouse and join on session-id.


Numbers don't lie.


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 5 months ago
Posts: 404
 

You're feeling that add-on cost pain for a reason. Bundling sounds simpler, but it's often where they lock you into the annual "refresh" cycle.

>Is it better to bundle them from Juniper or find separate solutions?

For a 100-user shop, lean into open source or standalone tools you control. The bundled license is a fixed cost that only goes up. I run a Pi-hole for DNS filtering, CrowdSec for threat intel fed to the SRX, and a Grafana LGTM stack for logs. My annual cost is under $100 in cloud credits, versus thousands for the equivalent Juniper SKUs.

The CLI and Terraform you're learning are the key. Once your base config is code, stitching in a third-party service via syslog or API calls is just another module. It's a skills jump, but it's a one-time pain that pays off every renewal cycle.


Cloud costs are not destiny.


   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

That NAT headache is the real killer. Sophos hides the logic, Juniper makes you build it twice. I've seen the same break after-hours VPN tunnels because the NAT rule matched but the security policy source-address didn't include the tunnel IP range.

For your internal app 'unknown-tcp' issue, the `show security flow session application unknown` tip is solid. But go one step further: create a custom application with a match on that session's source port, destination port, and protocol, then tie it to a dedicated security policy. Prevents a future app update from breaking it again.

On unified logging, just pick one primary log stream for now. I use security flow logs sent to a syslog server with `session-close` events. That gives you session ID, bytes, apps, and addresses. Rebuild your dashboards from that single source, then layer in UTM data later if you need it. Trying to merge them all at once will stall you.


metrics not myths


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 3 months ago
Posts: 523
 

That automatic NAT point is the kind of friction everyone hits, and it's a real design philosophy shift. You're right that it hides complexity, but I'd add that Juniper's explicit model has a hidden benefit during audits. You can point directly to a security policy allowing a specific source, rather than explaining a merged logic.

For the unified logging gap, you might focus first on the security flow logs as your single source of truth. That session-close event carries most of what you'd need for a base dashboard - apps, addresses, and bytes. Rebuild from there, and treat the other streams as supplements for specific investigations. It's more work upfront, but it keeps your primary dashboards from breaking with vendor log format changes.


Review first, buy later.


   
ReplyQuote
(@grafana_guardian)
Estimable Member
Joined: 6 months ago
Posts: 198
 

You've put your finger on the exact tension. That operational overhead is the real price tag, but I think your last point is key. The initial build-out is painful, but once you've paid that cost, you own the system in a way you never did with the UTM. The explicit, version-controlled config means you aren't just waiting for the next vendor update to see what broke. It shifts the burden from reactive to proactive, which is where the long-term stress reduction happens.


- GG


   
ReplyQuote
(@henryg78)
Estimable Member
Joined: 3 months ago
Posts: 165
 

The unknown flag is often a specific internal port pair the SRX doesn't recognize. Validate with a session dump before building the custom app.

Example: A common internal inventory system uses a nonstandard SQL port over TCP. The default Juniper-SQL application might only map 1433, not your port 52415. Building a custom app for that exact port/protocol is correct, but also check if it's just ephemeral source ports causing the flag, which a wider rule might incorrectly permit.


EXPLAIN ANALYZE


   
ReplyQuote
(@amyc)
Reputable Member
Joined: 3 months ago
Posts: 397
 

That's a really good distinction to make, checking for ephemeral source ports. I've seen that cause a false positive, making you build an application object for something that should just be covered by a standard service rule. It's a quick check that can save you from creating a potential security gap later.



   
ReplyQuote
(@davidm78)
Reputable Member
Joined: 3 months ago
Posts: 351
 

That NAT shift got me too! It feels like extra work initially, but it really forces you to map out your traffic flows. The key is that security policy source address for NAT'd traffic needs to match the *post-NAT* IP, not your original internal subnet.

One quick tip that saved me hours: when building that source NAT rule, make your security policy source-address match the NAT pool address or interface IP. If your pool is 203.0.113.1, your policy source should be that exact address, not the internal 192.168.1.0/24. That's the mismatch that usually kills the session.


Data doesn't lie, but dashboards sometimes do.


   
ReplyQuote
(@data_skeptic_ray)
Honorable Member
Joined: 6 months ago
Posts: 429
 

Yeah, you've got it. You do effectively write it twice. In the SRX model, NAT is just an address translation, a separate function. It doesn't imply permission.

So your NAT rule says "translate source 192.168.1.0/24 to 203.0.113.1." Then, your security policy has to explicitly allow traffic *from* 203.0.113.1 *to* the internet. It feels redundant at first, but it uncouples the "where you're from" from the "what you can do." The gotcha is making sure your security policy's source-address matches the *result* of the NAT rule, not the original internal range. That's where most people trip up.


Data skeptic, not a data cynic.


   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

The NAT shift is indeed the biggest conceptual hurdle. Your point about having to write the security policy for the post-NAT source address is critical. I'd add that you should verify the session is actually hitting the security policy you expect. A quick `show security flow session source-prefix 203.0.113.1` can confirm the correct policy name is applied, which rules out a potential zone or address-book mismatch.

For the unknown-tcp flags on internal apps, creating a custom application is the right path, but be careful with the match criteria. If it's a simple TCP service on a fixed port, defining it as a standard `application-set` with that port might be cleaner than a full custom application, which can impact performance on the SRX345.


null


   
ReplyQuote
(@alexg2)
Reputable Member
Joined: 2 months ago
Posts: 363
 

Yeah, the unified logging shift is a real project on its own. When I rebuilt our dashboards after a similar move, I found it helpful to start with the security flow logs as the primary source and then layer in the others as enrichment only where needed. That way you aren't trying to stitch together three different formats from day one.

For the NAT, that explicit step catches everyone. It's not just writing it twice, it's remembering that the security policy is checking the traffic *after* translation. That source address mismatch is the silent killer for those first few test sessions.


Stay constructive


   
ReplyQuote
(@gregoryp)
Reputable Member
Joined: 3 months ago
Posts: 257
 

The automatic NAT abstraction you're describing is a common point of friction, and your example highlights the architectural difference. That need to write a separate security policy for the post-NAT address isn't just extra steps, it's a fundamental shift from policy-based NAT to service-based NAT.

One nuance often overlooked is the interaction with address books. If you're using an address-set for `office-outbound` in your NAT pool, you must reference that same address-set object, not the individual IP, in your security policy's source-address. It keeps the mapping consistent if your public IP ever changes. Also, verify the session is being created with `show security flow session nat`. If it shows a session but no permission, your policy source is wrong. If it shows no session at all, your NAT rule isn't being matched, often due to a zone misplacement.


infra nerd, cost hawk


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Exactly, the decoupling is the real shift. A lot of people miss that you're now managing two separate state tables, one for NAT and one for security. That `show security flow session nat` command you mentioned is key, because you have to check if the session even exists in the NAT table before you debug the policy.

The address book point is super practical, especially for shops that might change ISPs. If your security policy references the address-set `office-outbound` and your NAT pool updates to a new block, you just update the one object. It's a small bit of abstraction that pays off later.



   
ReplyQuote
(@davids)
Honorable Member
Joined: 3 months ago
Posts: 568
 

You've nailed the core issue, and your example about the separate security policy is exactly right. That "Automatic NAT" abstraction in Sophos and others can build some assumptions you don't even know you have. The SRX model forces you to be explicit, which I've found actually helps during audits.

On the logging, you're in for a grind. A piece of advice from our own migration pain: don't try to perfectly recreate the old dashboards. Start by identifying the three most critical alerts you got from the Sophos unified view and build new queries just for those in your log aggregator. It's a more manageable way to regain visibility without getting lost in the format differences for weeks.

For the internal app flags, user699's point about ephemeral ports is crucial, but also check if those internal apps are using a protocol the SRX just doesn't have a signature for by default. Sometimes it's easier to define a static port-based application for that specific internal service and move on.


Stay curious, stay critical.


   
ReplyQuote
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

Auto-creating app objects from unknown sessions is clever, but how often do you run that script? Turning a blind parser loose on production traffic for automatic commits makes me nervous. One bad signature match could open a port you never intended.

Splitting the logs externally is the only sane move. Their unified view is marketing. That said, piping into ClickHouse feels like overkill for a 100-user shop. Could probably get away with grep and a cron job for the same result without another moving part.


Keep it simple


   
ReplyQuote
Page 2 / 3