Skip to content
Notifications
Clear all

Check out my Python script for auto-closing stale low-severity findings

38 Posts
37 Users
0 Reactions
83 Views
(@gracec)
Reputable Member
Joined: 3 months ago
Posts: 315
 

This is a classic scaling problem and your configurable thresholds are a good starting point. The dry-run is essential, but I'd caution that it can give a false sense of security if you're only looking at summary counts. I've seen scripts where a single misconfigured tag or an unexpected API filter change can cause a batch of incorrect closures that still looks reasonable in a summary report.

For a true safety net, your dry-run output should be detailed enough for a quick, but meaningful, human spot-check. Maybe it generates a random sample of 10 findings from the closure list, displaying the key attributes like asset tag, plugin ID, and age, so the reviewer can sanity-check the logic is catching the right things. It's about building trust in the automation before you let it run unattended.


The right tool saves a thousand meetings.


   
ReplyQuote
(@finnj)
Reputable Member
Joined: 3 months ago
Posts: 269
 

Oh, you've scripted your way into the classic compliance shortcut. Tag it as `environment:development`, wait 60 days, and poof, it's not a problem anymore.

But this just automates the risk acceptance, doesn't it? You're trading manual UI clicks for automated policy-based closure, but the core assumption is the same: old low-severity stuff is just noise. What happens when that dev server gets promoted to production by mistake and your script has already swept its history clean? You've erased the evidence of a potentially misconfigured asset.

A dry-run is nice, but it's just a preview of the same flawed logic. The real fix would be a script that *questions* why these findings are stale in the first place, not one that just hides them faster.


FOSS advocate


   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

Exactly, that dry-run summary is your first line of defense! I'd add that you should push it beyond just a count. Maybe have it spit out a few random samples of the actual findings it would close, so a reviewer can quickly spot-check if the logic is tagging the right assets. Helps catch those "why is this production server on the list?" moments before they happen.



   
ReplyQuote
(@gregr)
Reputable Member
Joined: 3 months ago
Posts: 343
 

Absolutely, the random sample is critical for building operator trust. I'd take it a step further and suggest the dry-run should emit two distinct lists: the random sample for a quick visual check, and a separate list of any findings that sit on the boundary of your filters, like assets that were tagged as 'production' in the last 24 hours or findings whose age is within a day of your threshold. That's often where the logic gets fuzzy and automation goes wrong.

You could also log the sample to the same run log that gets saved, so there's a permanent record of what the operator saw when they approved the batch. Otherwise, you're just trusting a console output that disappears.


throughput first


   
ReplyQuote
(@ide_tinkerer)
Reputable Member
Joined: 6 months ago
Posts: 338
 

Yeah, the ticket ID in the comment is a clever audit trail move. It turns the script's action from a vague "automated closure" into a specific, traceable business process step.

I'd be careful about relying solely on that comment field though. As someone else pointed out, it's often a tiny character limit. If your ticket ID is long or you need to add a reason code, you might hit the wall and truncate the audit trail you're trying to create. Maybe the script should first check the max length for the comment field via the API before it tries to write?

Also, what if the comment POST fails after the closure succeeds? Now you've closed something with no trace. The operation needs to be atomic, or at least have a retry loop for the metadata part.


editor is my home


   
ReplyQuote
(@davek)
Reputable Member
Joined: 2 months ago
Posts: 281
 

You're spot on about the API query being the bottleneck. In my tests, paginating through the findings took 80% of the runtime, even with parallel requests. The closure API calls are trivial by comparison.

> make sure your age filter uses the finding *last seen* date
This is critical. I'd also add a check for any state change in the last `n` days. If a finding was reopened or commented on recently, even if old, it should probably be excluded from the automated sweep. That's a common oversight in these scripts.


CPU cycles matter


   
ReplyQuote
(@devops_barbarian)
Honorable Member
Joined: 5 months ago
Posts: 439
 

Age-based cleanup assumes old findings are low risk. That's a dangerous assumption. Age is irrelevant if the underlying vulnerability never got fixed. You're not fixing the root cause, you're just hiding the symptom after an arbitrary timer runs out.

Your script might sweep a persistent low-severity config drift under the rug because it's "old." Now you've lost the ability to trend that issue over time.


Don't panic, have a rollback plan.


   
ReplyQuote
(@andrewb)
Reputable Member
Joined: 3 months ago
Posts: 292
 

You're automating the risk acceptance, not solving the problem. The real operational inefficiency isn't closing tickets, it's having a policy that generates meaningless findings in the first place. Your script just makes the dashboard look clean faster.

You'll miss the systemic issues because you're filtering them out. Congrats, you've built a tool that makes negligence scalable.


—aB


   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

Love the idea of logging the sample to the run log. That audit trail is essential. The console is ephemeral, and if you're ever audited or need to debug why a specific finding was closed, that saved sample is gold.

The boundary list is smart too. I'd expand it to include findings that have changed severity recently, like a low that was a medium last week. Those are always weird edge cases that deserve a second look.

One thing I've done is also tag those boundary items in the dry-run output with a specific flag, like [BOUNDARY], so they jump out during the visual scan. Makes the operator's job way easier.


✌️


   
ReplyQuote
(@harukik)
Honorable Member
Joined: 3 months ago
Posts: 400
 

Oh, I get why you'd want to automate that noise. It's a pain.

Quick question though - how do you handle the initial setup? Like, how do you decide on the specific age threshold or severity levels for your org? Is it just a best guess, or is there a process you follow to make those calls?

I'm looking at doing something similar but I'm stuck on the "what's the right number?" part.



   
ReplyQuote
(@gracehopper2)
Reputable Member
Joined: 3 months ago
Posts: 388
 

That's the hardest part, honestly. Starting with a best guess is fine, but you need to treat your first runs like a science experiment.

I'd pull historical data for closed low-severity findings and chart their age at closure. You'll probably see a cluster after 30, 60, or 90 days. That natural "cliff" is a better starting point than picking a round number. The goal is to codify existing team behavior, not invent new policy.

Just make your script's threshold a config flag from day one. Run it with a very conservative number first, like 180 days, and review every batch. You'll get a feel for what's truly "stale" in your context within a few cycles. Then you can tighten it down.


ship early, test often


   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

You're right to flag the API call multiplier. It's not just rate limits, either. Each extra round trip adds latency and multiplies the points of failure. If the closure succeeds but the comment fails, you're left in an inconsistent state, which is worse than no automation at all.

I've had to build wrappers around similar APIs for exactly this reason. The pattern is to treat the entire operation--fetch, decision, closure with metadata--as a single, idempotent unit of work with its own retry logic. You'd check for a `close_with_comment` endpoint first, and if it doesn't exist, implement a compensating transaction pattern: log the full intended state change locally *before* any API call, so if the comment fails, you have the context to retry it or at least flag the inconsistency.

The overhead isn't trivial. For a batch of 1000 findings, doubling the calls can turn a 2-minute job into a 4-minute one, and that's before you hit retries.



   
ReplyQuote
(@infra_architect_rebel)
Honorable Member
Joined: 5 months ago
Posts: 544
 

Compensating transaction is just adding more code to a problem you shouldn't have.

> overhead isn't trivial

Exactly. This whole script is overhead. The root problem is a tool generating meaningless noise. Adding retry logic and atomicity wrappers is polishing a turd.

Fix the policy. Stop generating the finding. That's the idempotent unit of work.


Simplicity is the ultimate sophistication


   
ReplyQuote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

Spot-checking random samples is smart, but I'd take it a step further. For any automated closure run, our team requires the output to be captured in the immutable run log itself, alongside the raw API query that generated the candidate list. That way, if someone asks six months later why a specific finding was closed, you can reconstruct the exact logic and dataset that triggered it. The visual scan is for the operator in the moment, but the forensic log is for the auditor later.

In practice, we include about five samples, but we also intentionally include the first and last items from the sorted result set. Often, edge cases live at the boundaries of your sort order (like oldest date or lowest severity score), and seeing those can reveal logic flaws the random sample might miss.


Logs don't lie.


   
ReplyQuote
(@infra_architect_rebel)
Honorable Member
Joined: 5 months ago
Posts: 544
 

Benchmarking the API is missing the point. The bottleneck isn't a tech problem.

You're optimizing a query for data you shouldn't need. If pulling findings is slow, you have too many to begin with.


Simplicity is the ultimate sophistication


   
ReplyQuote
Page 2 / 3