Skip to content
Notifications
Clear all

Just built a custom risk-based alerting workflow. My false positives dropped 60%.

5 Posts
5 Users
0 Reactions
13 Views
(@devops_dad_joke)
Reputable Member
Joined: 7 months ago
Posts: 288
Topic starter   [#26078]

Alright, let me tell you about my latest weekend project that finally made my on-call rotation bearable again. We were drowning in alerts from Splunk ES. Every little blip was triggering a major incident ticket, and the team was getting major alert fatigue. The classic "crying wolf" scenario, but with dashboards.

The core problem? Our ES correlation searches were running on a rigid schedule, treating a weird login from an employee's home IP with the same severity as a potential data exfiltration. My "aha" moment was realizing we had all this context *already in Splunk*—asset importance, user roles, geolocation—but we weren't using it to *triage the alerts before they became incidents*.

So I built a risk-scoring layer in front of our alert actions. Now, a raw detection fires, but instead of immediately creating a notable, it kicks off a lightweight risk assessment.

Here's the gist of the logic (simplified for clarity):

```python
# This is a conceptual outline of our risk engine script, fed by ES and lookup tables
def calculate_risk(event):
base_severity = event['severity']
user_risk = get_user_risk_score(event['user'])
asset_criticality = get_asset_criticality(event['dest'])

# Apply multipliers
if is_off_hours(event['time']):
base_severity *= 1.2
if geolocation_mismatch(event['src_ip'], event['user']):
user_risk *= 2.0

total_score = (base_severity + user_risk) * asset_criticality

# Decision tree
if total_score >= 80:
take_action('create_notable', event)
elif total_score >= 50:
take_action('send_to_slack_review_channel', event)
else:
take_action('log_only_to_siem_audit', event)
```

The key was enriching the raw event with our internal lookup tables (user roles, critical asset lists) *before* the risk logic runs. We used Splunk's `lookup` commands heavily in the initial search.

The result? A 60% drop in false-positive notables landing in our incident queue. The high-fidelity alerts now get immediate attention, and the medium-risk stuff goes to a dedicated review channel for triage without waking anyone up at 2 AM. The low-risk stuff just gets logged for compliance.

It's not perfect—tuning the risk weights is a continuous process—but my pager has been suspiciously quiet. Anyone else tried moving from a binary alerting model to a risk-based one? Curious about how you handle the risk scoring logic without making it a full-time job to maintain.

- tm



   
Quote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

That's a really smart approach, thinking about triage before the incident is even created. I've seen a similar problem in our ERP's audit logs where a failed login from a new region creates the same kind of panic as a system-level configuration change by an unauthorized user.

I'm curious about your asset and user risk lookup tables. Did you build those as static CSV lookups in Splunk, or are you pulling that context dynamically from another system, like a CMDB or an HR directory? I ask because we've tried to implement a similar scoring logic for inventory management alerts, but keeping the asset importance data current was a huge challenge. A server might be critical one day and decommissioned the next, and if the lookup isn't updated, your scoring is wrong.

How do you handle the risk assessment for events that involve multiple assets or users? For example, an unusual data transfer between two systems. Does the script take the higher of the two asset criticality scores, or average them somehow?



   
ReplyQuote
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 372
 

The static CSV lookup trap is a classic one, and the main reason these clever projects fail after six months. You're right to be wary. Pulling from the CMDB is the only way, even if it's messy. I use a Python script that runs as a modular input every five minutes, queries our ServiceNow CMDB API for the current `criticality_score` and `owner_department` fields, and dumps it into a KV store. The script costs nothing to run and it means a decommissioned asset's score drops to zero automatically.

For your multi-asset question, you're thinking about it backwards. Taking the higher or averaging assumes both endpoints contribute equally to the risk, which they don't. An unusual transfer *from* a HR database *to* a developer's test box is high-risk. The reverse direction is probably low. My scoring layer uses the source asset criticality as a multiplier and the destination as a modifier. So it's not symmetric, and it shouldn't be.


pay for what you use, not what you reserve


   
ReplyQuote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
 

I agree that CMDB integration is non-negotiable for long-term viability, but a modular input scraping an API every five minutes can become a scaling bottleneck when you have thousands of assets. It also creates a secondary data sync problem you now have to monitor.

A more resilient pattern is to have the CMDB publish asset state changes to a message queue, like Kafka. Your risk-scoring service subscribes and updates its internal lookup cache on-demand. This removes the polling interval and ensures your scoring logic reacts to changes in near real-time without hammering the CMDB API.

Your asymmetric scoring logic for source/destination is the correct architectural choice. It mirrors zero-trust network principles where intent and data sensitivity define the risk, not just the endpoints. Have you considered adding a third factor for the data classification label itself, if that metadata is available in your logs? A transfer of public marketing material between two high-criticality systems shouldn't score the same as PII.


infrastructure is code


   
ReplyQuote
(@bluepine)
Trusted Member
Joined: 2 months ago
Posts: 79
 

That's a really clever use of existing context to filter noise. I've been looking at similar workflows in Zendesk for support ticket routing, where you can apply risk scores based on customer history before assigning to a senior agent.

I'm curious, how are you managing the transition period with your team? When you introduced this layer, did you run it in parallel with the old system for a while to build trust in the new risk scores?



   
ReplyQuote