We've been running Palo Alto NGFWs for a few years. The rule base grew organically and felt messy. I'm not a networking expert, but I work in project management and wanted to understand our security posture better.
I spent some time this weekend building a Python script to parse our config. It looks for common issues. The initial results were surprising. Here's a sample of what it found across three firewalls:
- 28 rules with 'any' in both source and destination.
- 15 services set to 'application-default' but with no corresponding application object.
- Over 70 rules flagged as likely unused (zero hit counts for over 6 months).
- Several rules with overly permissive logging settings creating noise.
The script is basic, but it highlighted areas where we can clean up and reduce complexity. Has anyone else done something similar? I'm curious about what specific metrics or pitfalls others track in their rule bases.
Those findings are exactly the kind of hidden tech debt that builds up. The 'any/any' rules are a classic, but the 'application-default' with no app object is a sneaky one - it often means a rule is silently passing traffic you might not intend.
For metrics, I'd add checking for rules with very broad source zones (like 'any' or 'untrust') that allow inbound management protocols. Also, looking for rules that are shadowed by more specific ones above them - they're dead weight but can be tricky to spot manually.
Have you considered correlating your script's output with actual traffic logs, even just a sample, to see if those 'unused' rules truly have zero relevance, or if they're just for rare events?
Show me the accuracy numbers.
Parsing the config is a start, but it's a static snapshot. You're missing the runtime behavior. Zero hits for 6 months doesn't mean a rule is unused, it might mean it's a necessary emergency or break-glass path.
You should be more skeptical of your own script's logic. 'Any/any' rules are often intentionally placed at the bottom as an explicit deny-and-log catch-all. Blindly flagging those as an issue shows a lack of context.
How are you validating the findings? You'll cause an outage if you delete a rule based on a script you wrote over a weekend.
Don't panic, have a rollback plan.