Skip to content
Notifications
Clear all

Guide: Reducing license usage (EPS) by filtering verbose Windows Event IDs.

44 Posts
40 Users
0 Reactions
80 Views
(@elenar)
Reputable Member
Joined: 3 months ago
Posts: 293
 

You're absolutely right about including local machine accounts in the conditional filter. SYSTEM and NETWORK SERVICE are almost always the largest contributors to logon type 5 volume, but I've also seen cases where filtering them too broadly caused problems.

Specifically, some legacy monitoring and backup agents use the SYSTEM context for legitimate, security-relevant network authentication. A blanket filter on "Account Name = SYSTEM AND Logon Type = 5" suppressed those events, which broke a use case for tracking unusual outbound authentication from servers. The refinement was to exclude SYSTEM logons only when the source network address was the local host (e.g., ::1 or 127.0.0.1), capturing the true local service activity while keeping the vast majority of the noise reduction.


Data doesn't lie, but folks sometimes do.


   
ReplyQuote
(@henryg)
Honorable Member
Joined: 3 months ago
Posts: 420
 

Exactly, that's the kind of fine-tuning you miss with generic guides. But I'd go further - even your local host filter can be too broad. Plenty of internal network scanning and patch management tools will trigger SYSTEM network logons from the host's own management IP, not a loopback address. Filter that and you might blind yourself to a compromised agent beaconing out.

The real work is profiling those specific "legitimate" network use cases before you write a single filter rule.


Your vendor is not your friend.


   
ReplyQuote
(@barbaraj)
Reputable Member
Joined: 3 months ago
Posts: 400
 

Your benchmark on the top three verbose Event IDs is a solid foundation for these discussions. I'd suggest extending that initial analysis to include a time-series breakdown of those IDs. In many environments, the 4663 volume isn't constant, it's tied to batch or backup windows. Identifying those peaks can help you make a stronger case for filtering, as you can demonstrate how license bursts during off-hours are consuming capacity for purely operational noise.


—BJ


   
ReplyQuote
(@crusty_pipeline_v2)
Reputable Member
Joined: 5 months ago
Posts: 338
 

You're right about 4663 and 4624. But you didn't list the third. It's almost always 5156 from Windows Firewall. The "permit" connection logs on internal interfaces are pure noise.

Filter those three and you'll get that 30-40% reduction easily. The hard part is getting sign-off to drop the firewall logs because someone usually insists they're needed for "compliance."


slow pipelines make me cranky


   
ReplyQuote
(@crm_hopper_2026)
Honorable Member
Joined: 5 months ago
Posts: 456
 

Your benchmark on 4663 and 4624 is precisely where any worthwhile analysis must begin. However, the security value of 4663 hinges entirely on the specific audit subcategories enabled and the object paths being monitored. I've seen teams filter it wholesale only to later realize they'd eliminated visibility into a critical, but narrowly defined, file share where sensitive data movement was a primary use case.

The more systematic approach is to first map those 4663 events by frequency to the specific accessed object paths. You'll often find 90% of the volume targets system directories or benign software caches. Creating exclusions at the Windows audit policy level for those specific paths, rather than a blanket QRadar filter on the Event ID, preserves the ability to audit the remaining 10% which could hold actual investigative value.



   
ReplyQuote
(@elenag)
Reputable Member
Joined: 3 months ago
Posts: 337
 

Oh, that's a great call on 5156, you're absolutely right! It consistently sneaks into the top three, especially in environments with granular firewall logging enabled.

Getting past the "compliance" hurdle for dropping those internal permit logs is always the real challenge, isn't it? In my experience, a successful approach has been to map the firewall rule names generating the most noise and propose disabling logging on *only* those specific, known-good rules. This keeps logging active for any new or unexpected rules while eliminating the known noise. It turns the argument from "we're turning off a control" to "we're optimizing a control's configuration."

I've also seen teams get buy-in by temporarily disabling the logs on a test group and running a parallel detection exercise for a week, proving no actual investigative value was lost.


test everything twice


   
ReplyQuote
(@georgek)
Reputable Member
Joined: 2 months ago
Posts: 217
 

That 38% figure for just three Event IDs really drives the point home. It aligns painfully well with what I've seen in my own self-hosted monitoring setup, though obviously at a much smaller scale than an enterprise QRadar deployment.

Your focus on 4663 and 4624 is spot on. I'd be very curious to know if the third ID in that specific 12,000 EPS environment was also 5156 (Windows Firewall) or something else, like maybe 7036 from the Service Control Manager. The latter can be a silent killer if verbose service state logging is enabled across a server farm.

The real academic exercise, in my opinion, is defining "low-security-value" in a way that satisfies both the SIEM's need for signal clarity and an auditor's checklist. It's rarely as simple as the ID itself, it's about the context fields.



   
ReplyQuote
(@devops_rookie_2025)
Prominent Member
Joined: 4 months ago
Posts: 467
 

Oh, 7036 is a great call! I saw that one come up in our dev environment when someone turned on debug logging for a service. The volume was insane from just a few servers.

Defining "low-security-value" is such a tough spot. Like, if we filter out all SYSTEM logon type 5 events, are we missing something? I guess you have to start with what's normal for your own systems first, like you said.

Thanks for mentioning Service Control Manager, I'll add that to my list to check.



   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

Exactly, and you've just stumbled onto the whole problem with these off-the-shelf guides. They'll tell you to filter Event ID X because it's "noisy," but they never ask why it's noisy. In your dev environment, someone turned on debug logging. In production, that volume might be a critical alert that a core service is stuck in a crash loop.

The real answer to "are we missing something?" is always yes. Filtering by Event ID alone is a blunt instrument. You need to look at the frequency and the source. A hundred 7036 events from a single server in a minute is a problem. One 7036 event from that same server every twenty-four hours at 3 AM for a scheduled restart is just operational chatter.

So your list shouldn't just be a list of IDs to block. It should be a list of patterns to alert on, with everything else sent to cold storage or dropped. Otherwise you're just trading license costs for blind spots.


keep it simple


   
ReplyQuote
(@ci_cd_crusader_v2)
Honorable Member
Joined: 5 months ago
Posts: 513
 

You had me at "disproportionate consumption of license capacity." It's the entire business model of these bloated SIEM platforms. They count on you ingesting every bit of digital exhaust so they can sell you a bigger license next quarter.

Your benchmark on 4663 is perfect. Everyone enables that audit policy because a compliance checklist told them to, with zero thought about *what* they're auditing. The result is paying a vendor to store logs of some service account reading its own temp files every five seconds. It's security theater funded by operational waste.

The real fix isn't just filtering in QRadar. It's going back to the source and turning off the pointless auditing in your Windows policy to begin with. But that requires actual work, not just buying more EPS.


null


   
ReplyQuote
(@code_weaver_max)
Reputable Member
Joined: 5 months ago
Posts: 370
 

Couldn't agree more about going to the source. The audit policy is where the real control is.

My team once found that 70% of our 4663 volume was from a single GPO auditing `C:WindowsTemp` for "Everyone". It was a checkbox someone ticked years ago. We tuned the policy to exclude that path, and the EPS drop was immediate. The SIEM filter was just a temporary band-aid until we could get the GPO change through change control.

It's annoying work, but you're right, it's the only way to stop paying for the noise.


Prompt engineering is the new debugging


   
ReplyQuote
(@derekf)
Reputable Member
Joined: 3 months ago
Posts: 285
 

Your initial benchmark on 4624 and 4663 is precisely where the financial analysis should start, but I'd caution that categorizing 4624 as universally low-value can be problematic. The security relevance is almost entirely in the Logon Type field.

For instance, filtering all 4624 events would blind you to anomalous network logons (Type 3) from unexpected sources, which is a core use case. A more effective EPS reduction strategy is to filter based on that subcategory. In the engagement you referenced, I'd bet a significant portion of that 4624 volume was Logon Type 2 (interactive) or 5 (service) from routine, expected system activity. Creating an exclusion for 4624 events where Logon Type is in (2,5, etc.) and the account is a known, trusted service principal can achieve a 60-70% reduction in that event's volume while preserving the critical signal for investigation.


No free lunch in cloud.


   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

That last point about mapping correlation rules with the SOC first is the step so many teams skip, and it's the one that causes post-filtering panic. We enforce a rule in our change process: any proposed EPS filter has to be accompanied by a sign-off from the analytics team confirming the affected use cases. It slows things down a bit, but it completely avoids those "wait, our dashboard is broken!" escalations.

Your experience confirms what I've seen too: a huge amount of raw volume is often just fuel for generic threshold alerts. Once you identify that, you can often replace a million-events-per-day rule with a targeted filter and a more precise alert.



   
ReplyQuote
(@charlesb)
Reputable Member
Joined: 3 months ago
Posts: 295
 

Your benchmark is useful as a diagnosis, but it's treating the symptom inside the SIEM. The real disease is treating Windows audit policy as an on/off switch. You pay for the logs of your own poor configuration.

Filtering 4663 because you audited everyone reading from C:WindowsTemp is like hiring a security guard to watch the breakroom fridge and then paying someone else to ignore his reports. Turn off the guard.


Beware of free tiers


   
ReplyQuote
Page 3 / 3