Skip to content
Notifications
Clear all

How do I filter out low-priority alerts to reduce noise?

50 Posts
49 Users
0 Reactions
132 Views
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

>concrete, step-by-step approaches rather than just high-level ideas

Good, because the high-level ideas are usually wrong. The most effective step in LogRhythm isn't a rule parameter, it's turning rules off entirely. Start with a quarterly audit of rule efficacy. For each rule, pull a report of its offenses over the last 90 days and calculate the true-positive-to-action ratio. If a rule has fired 10,000 times and resulted in exactly zero documented incident response actions, you don't tune it, you disable it. The coverage is illusory.

Regarding rule logic versus thresholds, you must modify the logic. Threshold tuning is a reactive, endless loop. For a step-by-step example, take the ubiquitous "Windows Security Log Cleared" rule. Instead of alerting on every occurrence, the logic should condition on the absence of a preceding event ID 1100 (Audit Log Cleared Successfully) from the same source machine within the last 24 hours. This segments legitimate, logged maintenance from potential adversarial clearing.

On the AI Engine, treat it as a tertiary source of context, never a gatekeeper. Its risk score should be appended as a metadata field, like `ai_risk_score: 0.2`. Then, build a secondary correlation rule that requires a high traditional severity AND a high AI risk score to escalate to a Tier 1 queue. This forces consensus and prevents the AI's biases from becoming your blind spots. The Watchlist function is best used for suppression, not inclusion. Maintain a dynamic watchlist of known benign sources, like your patch management servers, and reference it as a "NOT ON" condition in noisy rules.



   
ReplyQuote
(@emilyr)
Reputable Member
Joined: 3 months ago
Posts: 295
 

I strongly agree with the quarterly audit approach, but the "true-positive-to-action ratio" metric needs a precise definition to be actionable. In my practice, we define an "action" as any incident ticket that transitions out of a 'New' or 'Triaged' state. This excludes automatic closes from noise-reduction workflows. We also plot this ratio over time; a rule with a stable 0.1% action ratio is a candidate for sunsetting, but a rule where that ratio is decaying from 2% to 0.2% over a year suggests the threat landscape or your environment has changed, and the logic needs a fundamental rewrite, not just a disable.

Your example of conditioning "Windows Security Log Cleared" on the absence of event 1100 is technically correct, but in modern environments, many legitimate log clears are scripted and may not generate that specific success event. We've had better results by joining the clearance event against a CMDB tag like `scheduled_maintenance_window: true` or an external orchestration system's API. The principle of adding context is key, but the source of that context often lies outside the SIEM's native log stream.

On the AI metadata point, appending the score is necessary but not sufficient. You must also log the specific factors that contributed to a low confidence score. We append a field like `ai_risk_factors: ["source_ip_internal_range", "time_business_hours"]`. This creates an audit trail that allows you to later validate if the AI's reasoning is sound, or if it's systematically downgrading a specific, legitimate pattern you've missed. Without that transparency, you're just trusting a different black box.



   
ReplyQuote
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
 

The decaying action ratio is the canary in the coal mine. If a rule's efficacy drops over time, it's usually because the ops team has automated the response you never documented. They've silently worked around the noise.

>source of that context often lies outside the SIEM's native log stream.

Exactly. That's where most filters fail. Pulling context from the CMDB or orchestration API is the only way. I've seen teams burn months tuning SIEM rules when the fix was a five-line Python script to fetch a maintenance calendar and stamp logs.


Prove it.


   
ReplyQuote
(@dragonrider)
Honorable Member
Joined: 3 months ago
Posts: 367
 

Great analogy with the email list segmentation, that's exactly the right mindset. Since you're asking for concrete steps, I'll give you the exact filter we built that cut our low-priority volume by about 40%.

For the "Failed Logon" rule, everyone's right about excluding patch system IPs. But the real step they're missing is layering in user context. We added a condition to *only* trigger the alert if the failing account has successfully logged into *any* system in the last 30 days. This filters out dormant service accounts, old test users, and attacker reconnaissance against dead accounts, which was a huge chunk of our noise. You can pull that active-user list from your HR system or even a simple daily Active Directory dump.

On the AI Engine, I use it as a post-rule triage flag, not a primary filter. If an alert fires from our tuned rule, the AIE appends a 'likely_context' field - like 'scheduled_task' or 'user_on_pto'. My team can see that metadata instantly and dismiss the alert in one click. It doesn't stop the signal, but it makes processing the noise 90% faster.

Have you looked at the source of your most frequent low-severity offenses? I bet a handful of systems or user groups cause 80% of it.


Try everything, keep what works.


   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

The user context layer you added to the failed logon rule sounds really smart. I'm just starting with our SIEM's alert rules, and that's a concrete step I can wrap my head around.

But how do you manage pulling that active user list? I'm trying to figure out the simplest way, like you mentioned an AD dump. Do you set up a scheduled task to export that list and feed it into LogRhythm daily, or is there a more direct connector?


Still learning


   
ReplyQuote
Page 4 / 4