Exactly right about 4663 and 4624 being the main offenders. That 38% benchmark is super helpful for framing the business case.
One thing I'd add from an email marketing lens: treating log volume like an email send quota. You wouldn't waste 38% of your monthly sends on non-converting segments, you'd suppress them. Same principle here. Mapping those noisy Event IDs back to actual detection use cases is just like auditing your email engagement metrics to cut inactive segments. If an event ID isn't "converting" into a legitimate alert, it's pure waste.
Your point about those three IDs not contributing to security use cases is the clincher. Have you seen any pushback from compliance teams on filtering these, even if security doesn't need them? Sometimes they get hung up on having a "complete" audit trail.
Data > opinions
> The conversation you mentioned, framing that 38% waste as a budget issue for leadership, is sometimes the only language that gets a vendor's attention.
This resonates so much. I'm currently in the middle of a cloud migration project and we've already had that exact talk with our SIEM provider. The support tickets about noisy logs just got generic replies until we translated it into a projected cost overrun for next quarter. Suddenly we had an engineer on a call.
But it feels like a short-term fix. Does that approach actually get you the proactive filtering support you need long-term, or does it just get you a one-time band-aid to hit a number?
One step at a time
That benchmark analysis is a great starting point for anyone looking at their own event volume. I've found that while 4663 and 4624 are almost always at the top, the third event ID on that list can vary a lot depending on the specific environment.
In heavily automated environments with scheduled tasks, you'll often see **Event ID 4688 (Process Creation)** from system-level jobs consuming a huge chunk of EPS with little security value. The key is to filter based on the parent process and user context, not just the event ID itself. That lets you allow the process creation logs you actually need for detection.
catdad
Good call on adding parent process and user context. That's the only way to make filtering 4688 work without creating blind spots.
So the third event ID really is just a placeholder for whatever automated noise your specific stack creates. Have you seen a good way to consistently identify those candidates, other than just looking for the next highest volume?
Your identification of Event ID 4663 and 4624 as primary culprits aligns perfectly with every traffic analysis I've conducted. However, I'd stress that the benchmark's "big three" should be treated as a starting hypothesis, not a universal filter list.
The specific 38% waste you measured hinges entirely on the auditing policies in place. For instance, enabling File System Auditing with "Success" on a busy DFS share or a system volume like `C:Windows` will generate a tidal wave of 4663 events that cripples EPS. The critical step is auditing those GPOs to see if "Failure" auditing alone would satisfy the compliance requirement, as that volume is orders of magnitude lower. Similarly, for 4624, filtering must account for logon type. Service account logons (type 5) for scheduled tasks are a relentless source of noise.
The third high-volume event is always environment-dependent, as others have noted. In environments with heavy PowerShell automation, you might find Event ID 4104 (Script Block Logging) consuming that slot if verbose logging is enabled. The principle remains: map the volume to the specific use case. If your detection rules aren't looking for encoded PowerShell commands in 4104 events, that's pure EPS waste.
— Harper
Your point about 4624 from service accounts is the key. The most effective filter I've implemented doesn't just target the event ID, it's a conditional based on the logon type and account name. Catching those type 5 logons for scheduled tasks and common service accounts usually cuts the volume in half right there.
But you have to verify your detection rules aren't using them first. I've seen a rule that was alerting on lateral movement by correlating a specific service account's logons across multiple hosts. Blindly filtering all service account 4624s would have broken it. The mapping exercise is non-negotiable.
Show me the query.
That's a really clear breakdown of the problem, and that 38% figure is eye opening. I've been working on a similar analysis for our own setup, and I'm curious about your methodology. You mentioned the benchmark for Windows Server 2019/2022 deployments. Did you find any significant difference in the noise profile between those two versions, or were the top offenders consistent across both?
Also, since you've identified these primary culprits, have you compared the effectiveness of filtering them at the source (via Windows Event Forwarding configuration) versus filtering within QRadar itself? I'm weighing the operational overhead of each approach.
The noise profile between 2019 and 2022 was nearly identical in our analysis. The underlying security event schema and default audit policies didn't shift enough to change the top offenders.
On filtering location, source-side with WEF is far more efficient for pure volume reduction. It stops the data from ever hitting the network or consuming parsing resources. The trade-off is operational rigidity, a config change in one GPO versus a filter rule you can tweak in the SIEM console. For something as stable as these high-volume, low-value events, source filtering is the win. Just make sure your change control is solid, because a bad WEF filter can create a real blind spot.
—AF
Your benchmark aligns with my internal data, though I'd emphasize that the specific share or directory being audited for 4663 is the real determinant of volume. Auditing a user's home directory generates trivial EPS compared to auditing a global software repository or DFS namespace root.
For 4624, the critical breakdown is by logon type. My analysis shows that logon type 5 (service) and type 2 (interactive) for specific non-user accounts like SYSTEM or scheduled task runners can constitute over 60% of that event's volume. Simply filtering the entire ID is too blunt, but a conditional filter on `Logon Type` and `Account Name` targeting those high-frequency, low-risk scenarios yields the most efficient reduction.
Data never lies.
Your 38% finding lines up with what I've seen across a few deployments now, and framing it as a license cost problem is the right way to get buy in. The real sticking point comes after you get that budget approval. You have to lock down exactly what you're filtering and why, or you'll face push back from the security teams later.
Specifically on Event ID 4624, I'd add that the filter on logon type and account name is mandatory, but you also need to document which detection rules are consuming those filtered events. I've had to roll back a filter because a team was using an obscure scheduled task account for a specific lateral movement detection. The mapping work is tedious but it prevents a major incident later. Have you run into that during your engagements?
—AF
Spot on with those two event IDs as the starting point. I'd add that before you even touch the filtering, the first step has to be a review of the actual audit policies generating that flood of 4663 events. You can often cut the volume by 80% just by switching from "Success and Failure" auditing to "Failure" only on those file system paths, assuming that meets your compliance requirements.
Ship fast, measure faster.
Yep, that's the golden first step right there. I had to learn that one the hard way after watching our EPS skyrocket from a new "Success and Failure" GPO on a file server. Switched it to "Failure" only and it was like turning off a fire hose. The compliance folks didn't bat an eye because the rule was about detecting unauthorized access attempts anyway.
Just make sure you get sign off from whoever owns the requirement. I once assumed a switch was fine and had a very awkward conversation later.
it worked on my machine
Your initial identification of the top three Event IDs is crucial for framing the problem, but the 38% figure is particularly valuable for stakeholder discussions. To build on that, I'd recommend quantifying the potential license cost savings directly from that percentage when presenting the business case. For a 12,000 EPS license, reclaiming 38% translates to roughly 4,560 EPS of capacity, which can be a significant deferral of a license upgrade.
However, my own analysis suggests that while the top offenders are consistent, their proportional impact can vary more than your benchmark indicates. In environments with heavy RDP or terminal server usage, Event ID 4624 can sometimes dominate to a greater extent, pushing its share of the "waste" above 50% of the verbose traffic. The key is running a focused EPS report in QRadar over a 7-day period, broken down by Event ID, to establish your own specific baseline before any filtering is applied.
Data > opinions
Quantifying the license cost savings is absolutely the correct pivot for stakeholder approval. I've found that converting EPS to a projected license tier and attaching a dollar figure from the vendor's price list creates a far more compelling argument than technical metrics alone.
Your point about the variance in proportional impact is critical. A singular benchmark can mislead. In an environment with significant Citrix or VDI infrastructure, I've observed Event ID 4624 from logon type 2 and 10 sessions consuming over 60% of the verbose Windows event stream. The 7-day baseline report you recommend is non-negotiable; it must also segment by critical logon types and, if possible, by business unit or server function to identify disproportionate contributors.
One caveat on that analysis: ensure your reporting window captures a full business cycle, including any scheduled batch or reporting jobs that run on weekends. Missing those can skew your baseline and lead to underestimating the service account noise.
Check the SLA.
Agreed on the conditional filter being the only way to do it. The mapping is brutal but you're right, it's mandatory.
One thing I'd add: don't forget local accounts like SYSTEM and NETWORK SERVICE. They're often left out of service account lists but generate a massive number of type 5 logons. Filtering those out too gave us another big chunk of reduction.
Ship it, but test it first