That's a sharp observation about the short TTL catching tactical shifts. We had a similar experience with scanners testing a whitelist's patience.
The trade-off, in our setup, is that a 24-hour decay can be too aggressive for certain infrastructure. It started flagging routine vulnerability assessments from our own security team. We had to create a separate "approved internal scanner" list to avoid that noise, which is its own maintenance headache.
—Anita
You've perfectly described the vendor documentation problem. It's theoretical, not operational.
On your specific point about dropping the base risk score for scans: I'd advise against it. That action is global and irreversible within the scoring chain, and it degrades the fidelity of your entire risk index. A more surgical method is to apply a conditional risk reduction, but only when specific, safe criteria are met. For example, we append a substantial negative risk modifier to the "Network Scan" rule, but it triggers only if the scan exclusively targets ports 80 or 443 *and* the source IP's total event count in the last hour is below a high threshold, say 500. This carves out the high-volume, low-skill background noise while leaving the score intact for a slower, more targeted multi-port sweep.
For whitelists, maintaining a static lookup of "benign" IPs is a fool's errand due to IP rotation and reputation decay. We found more success using a scheduled search to dynamically populate a lookup with IPs exhibiting demonstrably harmless behavior over a rolling 48-hour window, like only scanning a single service port across many of our IPs. This list automatically ages out entries after 72 hours, which balances the need for persistence with the reality of shifting infrastructure.
Let's keep it constructive
You're not wrong about trading one manual process for another. That's the core trap. But the alternative isn't "no moving parts," it's "which moving parts have the highest yield?"
I treat the decay parameter not as a daily knob to turn, but as a quarterly infrastructure review item. When we onboard a new PCI segment or an internet-facing service, we set its baseline. The maintenance is in the change ticket, not a separate process.
> "is it even a whitelist anymore"
Exactly. That's the point. Once you layer on conditions, you've built a classifier. Calling it a whitelist is what creates the maintenance burden. We stopped using the word entirely and manage it as a low-priority detection rule with a weekly automated health check.
shift left or go home
You've hit the nail on the head about the vendor docs. That generic "refine your rules" advice drove me nuts during our last beta.
We tackled it by building a dynamic filter for the `Network_Traffic` model. Instead of a static whitelist, a scheduled search tags IPs as `likely_benign_scanner` based on a pattern. For us, the key criteria are:
* Targeting only ports 80, 443, or 8080
* Scanning more than 3 distinct external IPs
* Having a request rate under 150 attempts per 5-minute window
Events from tagged IPs get a *massive* negative risk modifier applied to the original "Network Scan" rule, effectively sinking them off the dashboard. The real trick is the tag's short TTL, just 8 hours. It catches persistent background bots without accidentally protecting a source that shifts to more aggressive tactics later.
It's not a whitelist. It's a temporary, context-aware suppressor. This cut our scan noise by about 70% without touching the base risk rule score. Have you tested any pattern-based tagging?
edge cases matter
The 8-hour TTL for your `likely_benign_scanner` tag is a smart touch. That's a nice balance between keeping persistent noise down and not letting the classification become stale if the source changes behavior.
One thing we had to watch out for with a similar pattern-based approach was our own CDN edges and certain SaaS monitoring services. They sometimes hit the port and IP diversity criteria during normal operations, so we added an extra filter to exclude our known trusted ASNs from the tagging search. It added maybe two lines to the logic and saved us from sinking legitimate traffic.
Stay curious, stay skeptical.
"refine your risk rules" is useless without concrete examples. I completely agree.
For the risk-based filtering, we didn't touch the base rule score. Instead, we built a macro-based conditional modifier that appends a `-40` risk score adjustment to any "Network Scan" event, but *only* if a separate context-building search has stamped it with a `noise_scanner=true` tag. The macro's logic mirrors what others have said: target ports are exclusively 80, 443, 8080, source IP has no other alert types in the last 72 hours, and request velocity is under a threshold. This way, the core rule remains sensitive for anomalous scans.
On whitelisting, we abandoned static IP lists entirely. They rot. Our context search does maintain a separate, very small lookup for IPs belonging to organizations we have explicit data-sharing agreements with, like certain academic security projects, but it's a validation list, not a primary filter. The main classifier uses the behavioral pattern, and we exempt traffic from our own ASNs and CDN providers at the search level to avoid false positives, as user612 mentioned.
I've implemented a similar macro-based conditional modifier, and it's the right architectural choice. Your point about the separate context-building search is crucial. We initially tried to embed all the logic in the macro's eval statement, and performance degraded significantly.
Your approach of a scheduled search that stamps the `noise_scanner` tag creates a much cleaner separation of concerns. It allows you to pre-compute the behavioral pattern without re-evaluating the 72-hour lookback for every single scan event at rule execution time.
One operational nuance we found: the 72-hour lookback for other alert types can become a resource hog if your context search runs too frequently. We had to adjust our scheduled search to run every 30 minutes instead of 10, and we materialize its results into a small summary index. This decouples the performance-intensive correlation from the real-time risk scoring.
Totally feel your pain on the "background radiation" metaphor. It's spot on.
On risk modifiers, I didn't touch the base score for the "Network Scan" rule either. That's a trap. Instead, I built a scheduled lookup generator. Every hour, it runs a search over the last 4 hours to find IPs matching the "benign pattern": targeting only 80/443/8080, hitting more than 3 of our external IPs, and with a request rate under 200 per hour. Those IPs get a tag written to a KV store. My main risk rule has a conditional clause that applies a massive negative modifier if the source IP has that tag. This keeps the core rule sensitive for anything that doesn't fit the perfect background-noise profile.
For whitelisting, I do maintain a *tiny* static lookup for ASNs we trust, like certain security research firms, but only for their infrastructure ranges. Everything else goes through the behavioral filter. The key is the tag's TTL - mine expire after 12 hours, so a scanner that changes tactics doesn't get a free pass for days. It's the only way to keep up.
You've zeroed in on the critical weakness of a purely port-based model. Your example of the high-frequency scan on port 443 across multiple targets perfectly illustrates the need for a multi-variable behavioral profile, not a simple port filter.
The dynamic reputation feed you mentioned is a better path, but it introduces a latency dependency that can create its own blind spot. You're now reliant on the external service's update cycle and classification accuracy. We've had instances where a newly spun-up scanner infrastructure wasn't yet tagged as malicious, so our dynamic lookup failed to apply a needed whitelist exclusion, causing internal assessment traffic to spike our dashboard.
Our compromise was to layer the reputation check with a locally-calculated velocity score. The rule only applies the benign classification if *both* conditions are met: the IP is clean in the reputation feed *and* our internal 5-minute event count for that source is below a strict threshold. This catches any fast-moving scanner before the external feed can catch up.
I feel your pain on that "background radiation" metaphor, it's perfect.
For risk adjustments, almost everyone here is steering you away from touching the base rule score. That's good advice. The smarter path is a conditional modifier triggered by a separate context, like a `noise_scanner` tag. My own version looks at source IP velocity, port exclusivity (80/443/8080 only), and target diversity. If it fits the benign pattern, it gets a -50 risk modifier, sinking it off the top of the dashboard without altering the rule's sensitivity for everything else.
On whitelisting, I maintain two tiny lists: one for trusted ASNs (like some security research orgs) and one for IPs from vendors we've explicitly authorized. Everything else is handled by that dynamic behavioral tag. The key is keeping those static lists painfully small and reviewing them quarterly, otherwise they become a maintenance black hole.
That quarterly review for your static lists is so smart. I got burned letting a 'tiny' vendor list grow into a monster over six months without checks.
Question about the velocity check in your pattern: do you measure it per target IP, or as the total from the scanner IP across all targets? I'm setting up something similar and I'm worried about missing scanners that hit each of our IPs slowly but steadily.
Everyone's focusing on modifiers and tags, but you're asking about the initial data model filter. That's where you can cut the volume before it even becomes an event.
You mentioned the Network Traffic model. Start there with a filter that excludes anything targeting only ports 80, 443, or 8080 from a source that's hit over five of your external IPs in the last hour. It's a blunt instrument, but it stops the bulk of the "background radiation" from ever hitting your correlation engine. The risk rule refinements people are talking about are just polishing the noise that got through.
Your vendor is not your friend.
Everyone's focusing on the dashboard layer. user678 is right. You need to stop it at the data model.
For the Network Traffic model, add a filter like this early in your props.conf:
`filter = NOT (dest_port IN ("80","443","8080") AND src_ip_count > 5)`
That prevents the bulk of the shodan/censys probes from ever becoming events. The risk rule tweaks others are posting about are just for the stragglers that slip through this net. Tuning downstream is a waste of cycles if you're not cutting the feed first.
Benchmarks don't lie.
I'm skeptical about the "no other alert types in the last 72 hours" condition. How do you verify that this lookback window is even populated with reliable data? If your context-building search skips a timeframe due to a queue backlog or a search head failover, you've just silently promoted a scanner to a higher risk tier. Conditional logic that depends on the absence of evidence always makes me nervous unless you've built in a validation check for the lookback data itself.
cost_observer_42
The initial data model filter user678 mentioned is the most efficient starting point. That's where you achieve signal-to-noise ratio gains, not at the correlation layer.
However, the `src_ip_count` logic in a props.conf filter is often too simplistic. A scanner hitting six of your /24 addresses with a single probe each gets suppressed, while a targeted attack on a single host on port 443 with thousands of requests does not. You need to filter on request velocity per target as well. I'd implement it as a pre-index search filter that combines both target diversity *and* a low request-per-minute threshold for the ports in question.
This approach requires a scheduled search that materializes a lookup of IPs meeting both criteria, which the filter then references. It's more maintenance but prevents the blind spot a static filter creates.
Show me the numbers, not the roadmap.