Skip to content
Notifications
Clear all

Guide: Reducing ES index footprint by pruning unneeded data sources.

27 Posts
26 Users
0 Reactions
46 Views
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
Topic starter   [#25956]

Alright, gather 'round the digital campfire, folks. Been running Splunk ES for years now, and if there's one lesson my storage array has screamed at me, it's that we're data hoarders by default. Every sourcetype, every log source, gets vacuumed up and tossed into those precious ES indexes. But do you *really* need Windows Security events from that print server for your security analytics? Probably not.

The trick isn't just turning off data collection at the forwarder—though that's part of it. It's about being surgical *inside* Splunk. Let's talk about pruning at the index level. You can use `props.conf` and `transforms.conf` on your indexers or heavy forwarders to drop entire event streams you know are noise for ES.

Say you've got a fleet of network devices sending syslog, but only the firewall logs are relevant for your correlation searches. You can route all that data to a raw `net_raw` index, then use a nullQueue transform to only send the firewall stuff to your primary security index (`es_events`).

Here's a simplified example:

**props.conf**
```
[cisco:asa]
TRANSFORMS-filter = security_only
```

**transforms.conf**
```
[security_only]
REGEX = %ASA-(1-6)-d+
DEST_KEY = queue
FORMAT = indexQueue
```

For events that don't match that critical ASA severity (1-6), you'd route them elsewhere or drop them. The real power move is auditing your existing ES data. Run a search over the past 30 days to see what sourcetypes are actually being used by your correlation searches.

```sql
index=es_events earliest=-30d
| stats count by sourcetype
| lookup es_correlation_searches sourcetype AS sourcetype OUTPUT search_name
| where NOT isnull(search_name)
```

That'll show you what's actually feeding your detections. Anything not in that list? Prime candidate for the chopping block. Start by routing that data to a cheaper, non-ES index. Wait a week, make sure no alerts break, then you can stop collecting it entirely. My last cleanup at the day job reclaimed nearly 40% of our daily ES ingestion—that's real cash money back in the budget, and faster searches to boot.

Remember, kids: in security, more data isn't always better. *Relevant* data is. Now go check what that mysterious `sourcetype=api_audit` from that old dev box is actually costing you.

-- Dad


it worked on my machine


   
Quote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

Solid approach, and you're right about targeting at the indexer. One caveat: that nullQueue method on the transforms can backfire if your regex isn't airtight. A single misplaced character and you've silently dropped critical data.

I'd add that you should validate this with a test index first, or use a SEDCMD to comment the events before outright dropping them. It's too easy to create a blind spot.

Also, keep an eye on your data model acceleration summaries. If you prune source data they depend on, you'll see failed accelerations and that hurts ES performance more than the storage savings help.


—AF


   
ReplyQuote
(@emmam)
Estimable Member
Joined: 2 months ago
Posts: 216
 

Totally agree on the validation step. I've set up a pre-prod index specifically for testing these transforms, and it's saved us a few times.

Your point about data model accelerations is super important, it's a ripple effect that's easy to miss. One more thing to check: if you're pruning sources that feed notable event aggregations, make sure your correlation searches still have enough data to trigger correctly. A quiet search is its own kind of blind spot.



   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

That's a perfect example of the surgical approach. Just a heads up on the regex, though: you'll want to catch the full range of ASA severity levels. The example you gave might only match 1 through 6, but you'd typically want `%ASA-[0-7]-` to include everything from emergencies (0) to debugging (7). A small tweak, but it prevents unexpected events from slipping through.

Also, always pair this with a `nullQueue` destination for the filtered data to actually drop it. The DEST_KEY line is critical for that.


Keep it civil, keep it real


   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 3 months ago
Posts: 254
 

Pre-prod testing is the right call, but I'd shift that validation one step earlier in the pipeline. If you can move the filtering logic to a heavy forwarder tier, you can test the transforms on a small subset of data before it ever hits an index. This isolates the risk from your main indexing infrastructure.

The correlation search blind spot is a subtle but critical point. It's not just about having enough data, it's about preserving the event *distribution*. If you prune 95% of routine Windows login failures, the remaining 5% might be statistically anomalous enough to skew your baselines and actually cause *more* false positives, not fewer.


Data is the source of truth.


   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

Excellent catch on the regex range! That's the kind of detail that slips through and causes a headache six months later. It makes me wonder how many other default sourcetypes have similar little gotchas in their severity or priority fields that we might miss when building these filters.

You're absolutely right about the DEST_KEY being critical, too. I've seen configs where someone writes a perfect TRANSFORMS rule but forgets that line, so it's classifying the data but not actually dropping it, which defeats the whole purpose. The data just sits there, silently tagged, still consuming license and resources.


Pipeline is king.


   
ReplyQuote
(@calebh)
Reputable Member
Joined: 3 months ago
Posts: 421
 

Great example of index-level filtering. I'd add that you should also consider filtering by source or host, not just sourcetype like `cisco:asa`. A single sourcetype often comes from many devices, and you might only want to drop events from specific, low-value hosts like that print server you mentioned. A regex matching the host field in a transform can be just as effective for surgical removal.


Trust the data, not the demo.


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 3 months ago
Posts: 523
 

Absolutely right about filtering on host or source. It's often the more precise tool, especially when dealing with a diverse device fleet under one sourcetype. The trick is ensuring your host field is consistently and reliably populated; if forwarders are setting it, you need to trust that process.

One caveat here is that over-pruning by host can sometimes fragment your data model lookups if those hosts are referenced elsewhere. It's less common than the acceleration issue, but worth a quick check in the data model definitions before you commit.


Review first, buy later.


   
ReplyQuote
(@emilyh)
Estimable Member
Joined: 3 months ago
Posts: 166
 

That's a good point about DEST_KEY being easy to forget. It makes me think a simple validation step could be added to the deployment process, like a script that checks any transform with 'nullQueue' actually has that key set.

On the regex gotchas, is there a reliable way to audit a sourcetype for common field patterns like that, or is it mostly manual review of sample data?



   
ReplyQuote
(@danm)
Honorable Member
Joined: 3 months ago
Posts: 452
 

Spot on about host filtering being more precise. I've used it to exclude noisy dev servers from prod indexes, and it's way cleaner than sourcetype-level drops.

The biggest gotcha I've run into is host field inconsistency, especially when you have multiple forwarders or universal ones. One bad rename and your filter silently misses a chunk of data.



   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

Host filtering is precise until it isn't. The regex approach you mentioned works if your host field is static and perfectly normalized. In dynamic environments with auto-scaling groups or containers, that's a huge assumption.

Better to tag low-value data at the source with a universal metadata field, then filter on that. It's one extra step but decouples your filtering logic from volatile infrastructure identifiers.


Least privilege is not a suggestion.


   
ReplyQuote
(@daisym)
Reputable Member
Joined: 3 months ago
Posts: 226
 

Oh, that example hits home! I had to do something similar with our web server logs. Tons of health check pings and crawler traffic were bloating our main security index.

Just watch out for the acceleration rebuilds after you apply that transform. If your firewall data model is accelerated, pruning the raw stream like that can trigger a full rebuild of that summary, which can be a hefty load on the indexers for a bit. Might want to schedule the change during a maintenance window.



   
ReplyQuote
(@emilyk99)
Estimable Member
Joined: 2 months ago
Posts: 173
 

The nullQueue approach makes a lot of sense for isolating relevant traffic. When you set this up, how do you monitor that the filter isn't accidentally dropping something you need? I'm thinking about setting up a parallel index or a summary alert for when the filtered volume suddenly drops to zero, but I'm not sure if that's overkill.

Also, once you start filtering at this stage, does it affect how you calculate your license usage? I assume the data sent to nullQueue still counts as ingestion, since the indexer has already processed it.



   
ReplyQuote
(@cost_observer_42)
Honorable Member
Joined: 4 months ago
Posts: 407
 

Exactly. That host field is a house of cards. You're trusting a configuration spread across dozens or hundreds of forwarders to be perfectly consistent, forever. One team decides to change their host naming convention and your filter is suddenly just a suggestion.

And it's worse than just missing data. The data you *wanted* to drop is now in your prod index, costing you money and noise, but you won't see an ingestion drop to alert you. The billing looks the same. You only find out when your searches get slower and your license pool runs dry a week early 😒

Has anyone actually built a reliable audit for this, or are we all just crossing our fingers and checking the volume reports every Monday?


cost_observer_42


   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 3 months ago
Posts: 380
 

That's a solid technique for index-level pruning, but your regex example might be overly restrictive for real firewall logs. The `%ASA-(1-6)-d+` pattern only matches messages with a single-digit message ID following the dash, like `%ASA-6-305011`. Many useful ASA messages have multiple digits, like `%ASA-4-106023`.

A more comprehensive pattern would be `%ASA-(1-6)-d{5,6}` to capture the standard five or six digit message codes. Even better, you could invert the logic and filter out the known noise message ranges instead, assuming you're confident in what constitutes 'security' events.


null


   
ReplyQuote
Page 1 / 2