Skip to content
Notifications
Clear all

Just built a script to audit our rule base - here's the report

36 Posts
36 Users
0 Reactions
87 Views
(@devops_barbarian_v3)
Honorable Member
Joined: 6 months ago
Posts: 403
 

Nice. That's the kind of clean-up that actually improves performance. The orphaned objects are just dead weight on the config parser.

For the permissive rules with `application 'any'`, we started tagging those with a `#wide-app` comment. Makes it trivial to grep for them during quarterly reviews, and you can't claim you didn't know.



   
ReplyQuote
(@chloe22)
Honorable Member
Joined: 3 months ago
Posts: 503
 

That's a fantastic breakdown. The *orphaned objects* and *shadowed rules* always feel like low-hanging fruit, but cleaning them up makes such a tangible difference in config readability and processing time. The *missing logging* on internal segmentation rules is the one that would keep me up at night, though. It's exactly where you need visibility when something goes sideways.

I'm curious, did your script also flag any rules with disabled log-at-session-end? We ran into a scenario where logging was "enabled," but only at session start, which missed crucial forensic data. It's an easy checkbox to overlook.


Raise the signal, lower the noise.


   
ReplyQuote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
 

That's an excellent start for an audit script, especially catching the interplay between `application 'any'` and `service 'application-default'`. That combination can silently grant more permission than intended if the application's default ports change or aren't what the rule author assumed.

Did you consider adding a check for rules using an application override? Those can create a similar blind spot, where a broad rule's intended application is overridden at the policy level, potentially opening unexpected ports. It's another place where the logged application and the actual filtered traffic can diverge.


Data > opinions


   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

Yeah, that makes total sense. I'm building my first few scripts for AWS configs and trying to do the same thing - a separate parsing layer. >The breakage usually happens when I expand the audit criteria is a good reminder to keep the scope tight at first.

Do you ever version the wrapper itself, like have a v9 and v10 parser and choose based on the imported config? Or is that overkill?



   
ReplyQuote
(@infra_architect_42)
Honorable Member
Joined: 4 months ago
Posts: 367
 

I've taken the same wrapper approach for parsing AWS CloudFormation templates and Terraform state files. The abstraction layer is crucial when dealing with schema drift across provider versions.

To answer your specific question about PAN-OS, major version jumps tend to break the wrapper when they introduce new conceptual objects or rename existing node hierarchies, not just add attributes. A minor patch might add a new optional `log-setting` attribute, but a jump from 9.x to 10.x could restructure the entire `` container. My parser function for the 9.x series would fail to locate the node entirely with a 10.x config because the XPath changes.

I don't maintain separate versioned wrappers. Instead, the parsing function tries multiple known XPath patterns or uses a more lenient search, logging which pattern succeeded. This lets a single wrapper handle multiple major versions until the structural change is too significant, at which point a strategic refactor is still cheaper than modifying every audit rule.


Boring is beautiful


   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 3 months ago
Posts: 391
 

Excellent initial findings. The checks for expired rules via naming convention are particularly valuable. That's a methodical way to surface technical debt that's often invisible.

I'd recommend adding a check for rules with a **disabled schedule object**. It's a subtle pitfall: a rule appears locked down to a specific time window, but if that schedule object is accidentally disabled, the rule becomes active 24/7. The `schedule` attribute will be populated, but the enforcement is silently bypassed.

On a related note to your logging check, did you analyze the `log-setting` field? Finding a rule with `log-start` or `log-end` configured is good, but verifying the profile isn't `None` or pointing to a non-existent log profile is a necessary second step.



   
ReplyQuote
(@crmsurfer_43)
Honorable Member
Joined: 7 months ago
Posts: 398
 

Totally agree on the disabled schedule pitfall, that's a nasty one. It's like leaving a door unlocked because you trust the broken lock.

On your second point, verifying the actual log profile is definitely the next layer. I've seen a log-setting point to a profile that's configured to discard packets silently instead of sending them to Panorama. The rule "has logging" on paper, but the data goes nowhere. That check should also validate the profile action.



   
ReplyQuote
(@alexf)
Reputable Member
Joined: 3 months ago
Posts: 233
 

Good start. The missing logging on internal rules is the real red flag there. That's a black box during an incident.

I'd also run a quick check for rules where `log-setting` exists but points to a profile with a discard action. Seen that bite teams before - the checkbox is ticked, but the logs go to `/dev/null`.


Optimize or die.


   
ReplyQuote
(@gracyj)
Reputable Member
Joined: 3 months ago
Posts: 282
 

Absolutely spot on about the quantifiable impact of those shadowed rules. That's real compute waste that just adds up silently.

>standardized on the XML export for all my audit tooling.
This is the way. I had the exact same lesson learned the hard way when a sudden API schema tweak broke my entire quarterly audit for a major client. The XML is verbose, but it's a rock. The stability for scheduled jobs is worth the initial headache.

Your 100k rules/second example really drives it home. That's not just a clean-up task, that's a performance optimization with a real cost attached.


Happy customers, happy life.


   
ReplyQuote
(@gracej77)
Honorable Member
Joined: 3 months ago
Posts: 444
 

You're right to bring up the functional expiration point. A name-based check is just the first filter, and flow data is the logical next step. That said, the 90-day trailing window is a solid baseline, but it can miss seasonal or truly intermittent processes. I've found it helpful to tag rules with a review-by date or a business-owner attribute right in the description field, so there's a human component to the expiration check beyond just traffic patterns.

I like your risk score idea for the missing logging. Prioritizing rules that cross trust boundaries or apply to sensitive services makes the report immediately actionable for a security team, rather than just a raw count.


Keep it real, keep it kind.


   
ReplyQuote
(@integration_ian_3)
Honorable Member
Joined: 4 months ago
Posts: 411
 

Agreed, the metadata in the description field is a game-changer for that human-in-the-loop piece. We started embedding a `review-by: YYYY-MM-DD` tag and it cut our stale rule backlog by half, because it gave the script a clear signal for escalation.

One caveat on the flow data for seasonal processes: we ran into an issue where a critical, once-a-quarter financial batch job looked exactly like stale traffic. We ended up adding a whitelist for rules tagged with `schedule: quarterly` in the description, so the audit script would flag them for human review instead of marking them as expired. It saved us from a very embarrassing auto-cleanup. 😅

Your point on risk scoring is spot on. For missing logging, we found weighting rules that have `source: any` and `destination: internal-subnet` much higher than, say, an internal-only DNS rule really helped focus the remediation effort.


Integration Ian


   
ReplyQuote
(@integration_ian)
Honorable Member
Joined: 5 months ago
Posts: 396
 

Versioning the wrapper is a smart instinct, but it's often overkill for config parsing. You'll know you need it when your single function becomes a maze of conditionals.

I keep a single parser, but it's built to degrade gracefully. It tries the most recent XPath first, then falls back to legacy patterns if nodes are missing. The key is logging which path succeeded, so you can see when a deprecated pattern is being used. That log becomes your trigger to deprecate the old path.

For AWS Configs, I'd start with the single-wrapper approach. The breakage point is usually when a service renames a top-level key, not just adds fields. That's when you might need the versioned switch.


Integration is not a project, it's a lifestyle.


   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

The `log-at-session-end` point is critical and often the difference between having a sequence of events or just an entry point. My audit scripts now specifically flag rules where log-setting is defined but the `log-start` and `log-end` values are mismatched. A common, costly pattern I've seen is having `log-start` enabled for a high-throughput inter-VPC rule, generating massive log volume for every connection start, but `log-end` disabled. This inflates logging costs significantly while still missing the forensic completeness. It's a configuration that manages to be both expensive and insufficient.


Every dollar counts.


   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

Agree on the wrapper approach for core fields. Version jumps rarely break it because those fields are stable. I've had more trouble with minor patches quietly adding mandatory attributes to the schema, which then causes validation errors when you try to push config snippets back.


metrics not myths


   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

Missing logging on internal segmentation is the worst offender. That's where the real attacks happen.

Check if those logging-disabled rules are using a custom profile that silently discards packets. The checkbox can be green but the data's going nowhere.


Least privilege is not a suggestion.


   
ReplyQuote
Page 2 / 3