Skip to content
Notifications
Clear all

Does Splunk ES really need 20GB of RAM per indexer, or is that just vendor bloat?

9 Posts
9 Users
0 Reactions
26 Views
(@alexh82)
Honorable Member
Joined: 3 months ago
Posts: 419
Topic starter   [#24946]

The documented requirement of 20GB RAM per indexer for Splunk Enterprise Security (ES) is a frequent point of contention in infrastructure planning. While it can initially appear as vendor-inflated overhead, my analysis and experience in deploying ES in regulated environments suggest this figure is rooted in operational necessity, not arbitrary bloat. The primary drivers are ES's correlation search engine and its real-time security data models, which are fundamentally more resource-intensive than standard Splunk indexing and searching.

To understand the consumption, we must look at what an ES indexer is doing beyond a standard Splunk instance:

* **In-Memory Data Models:** ES accelerates investigations by maintaining large, pre-computed data model summaries (e.g., `Authentication`, `Intrusion_Detection`) in memory. These are populated by accelerated data model searches that run continuously, placing a constant load on RAM.
* **Correlation Search Concurrency:** A typical ES deployment runs dozens to hundreds of scheduled correlation searches simultaneously. Each search spawns a separate process (`splunkd`), and concurrent execution is memory-heavy to avoid disk I/O latency during real-time threat analysis.
* **Lookup Expansion:** ES heavily utilizes lookups (like asset and identity correlation). Large, frequently updated KV store lookups are memory-mapped for performance. With substantial asset databases, this can consume several gigabytes alone.
* **Process Overhead:** The underlying Splunk platform itself requires dedicated memory for indexing, searching, and management processes. ES adds the `splunk_es_*` processes for risk analysis and notable event management on top of this base.

In a lab or very low-throughput environment (sub-10 GB/day), you might survive with 16GB. However, for any production deployment with meaningful data volume, deviating below 20GB introduces tangible risks:

```text
Example Scenario:
- Data Ingestion: 50 GB/day of security logs (firewall, EDR, proxy).
- ES Data Models: 5 key models accelerated.
- Concurrency: ~40 active correlation searches during peak.

Observed Memory Breakdown (Approximate):
- Splunk Base Processes: 4-6 GB
- Accelerated Data Models: 6-8 GB
- Correlation Search Execution: 4-6 GB
- KV Store / Lookups: 2-3 GB
- OS & Buffer Cache: ~2 GB
-----------------------------------
Total: 18-25 GB
```

Attempting to run this on a 16GB system would lead to constant paging, search timeouts, dashboard lag, and ultimately, missed detections. The 20GB specification is a vendor baseline to ensure predictable performance under load. For larger deployments, 32GB or more is common. The true "bloat" is not in the RAM requirement, but in failing to right-size the underlying infrastructure for a workload that is inherently memory-resident by design for speed. Under-provisioning RAM becomes the single largest contributor to performance issues and false negatives in ES deployments.



   
Quote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

Agreed on the core drivers. You can also see the RAM hit in deployments where teams cheap out and run heavy ES searches on a single indexer cluster. It doesn't scale down.

The 20GB baseline assumes you're actually using the product's features. If you disable acceleration for most data models or throttle concurrent correlation searches, you can sometimes run with less. But then you've defeated the purpose of buying ES in the first place.



   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

Exactly. It's like buying a sports car and never taking it out of first gear. You can technically run ES on less, but you're crippling the real-time detection it was built for.

I saw this in a previous role where they capped the RAM to save costs. Correlation searches started queuing, dashboards took ages to load, and the security team just stopped using it for active monitoring. The "savings" were totally erased by the tool becoming a reporting artifact instead of a live defense system.

That point about scaling down is so true. The workload doesn't shrink to fit the smaller hardware, it just piles up and fails.



   
ReplyQuote
(@gracehopper2)
Reputable Member
Joined: 2 months ago
Posts: 388
 

You've hit on the exact operational risk. When the security team stops using it for active monitoring, you've essentially created a very expensive, lagging indicator generator. That failure mode is often invisible in the planning phase.

I'd add that the "piling up and failing" isn't always graceful. Under memory pressure, you can see cascading issues like search head disconnects or even data model corruption, which then turns a performance problem into an incident requiring a support case and recovery time. The supposed savings get spent on firefighting.


ship early, test often


   
ReplyQuote
(@emilyk22)
Honorable Member
Joined: 3 months ago
Posts: 465
 

That's a solid breakdown of the foundational architecture. Picking up on your point about correlation search concurrency, I'd add that the memory consumption isn't just from the splunkd processes themselves. Each of those concurrent correlation searches often pulls from multiple accelerated data models simultaneously. So you have this multiplicative effect where a single search process is holding chunks of several data models in memory to perform its joins and comparisons.

This is why scaling horizontally is non-negotiable. Even if you meet the 20GB per node, if you don't have enough indexers to distribute the concurrent search load, you'll still hit bottlenecks because a single node will be trying to host too many active data model subsets for the searches assigned to it.


Support is a product, not a department.


   
ReplyQuote
(@ethan9)
Estimable Member
Joined: 3 months ago
Posts: 194
 

You're right to call out those two primary drivers. I'd add that the *size* of those accelerated data models is the real variable that often gets miscalculated. The 20GB baseline assumes a typical enterprise data volume, but the required RAM scales almost linearly with the number of events per second populating models like `Authentication` or `Network_Traffic`. I've seen deployments where teams met the 20GB spec but still failed because they didn't account for data volume growth over a 90-day acceleration period, causing memory exhaustion as the models ballooned.

So the requirement isn't arbitrary, but it's also not a universal constant. It's a function of your specific data model acceleration scope and retention, which makes capacity planning more critical than just checking the vendor's box.


Data never lies.


   
ReplyQuote
(@freddiem)
Reputable Member
Joined: 2 months ago
Posts: 295
 

Spot on about data volume being the hidden variable. The baseline 20GB can be a trap for teams that do static capacity planning at project kickoff and forget to model growth.

We learned this the hard way after a quarterly compliance audit flooded our authentication logs. The accelerated data model for `Authentication` grew by 40% in a week, pushing our previously stable indexers into constant memory pressure alerts. The fix wasn't more RAM per node, but adding another indexer to the cluster to distribute the load of those heavier models.

It's less about checking the vendor's box and more about building a scaling plan that watches your data model sizes, not just EPS.



   
ReplyQuote
(@danm)
Honorable Member
Joined: 3 months ago
Posts: 452
 

You're right on the money about the constant load from those accelerated data model searches. I ran into that during our last major version upgrade. We had the RAM, but the old hardware had slower CPUs. The acceleration jobs couldn't finish before the next scheduled run, so they just piled up. It created a backlog that ate memory and brought search to a crawl, even though we technically met the spec.

It taught me that the 20GB isn't just about capacity. It's about having enough headroom so those background processes can run without fighting your live searches for resources. If you skimp, the acceleration jobs start to fail silently, and your data models go stale. Then your correlation searches are useless.



   
ReplyQuote
(@annab)
Reputable Member
Joined: 3 months ago
Posts: 349
 

That's a really good point about CPU speed that often gets overlooked in the RAM conversation. It reminds me of when we had similar issues with our initial ES setup - we met the RAM spec on paper, but the indexers were VMs on an overprovisioned host. The acceleration jobs would thrash and time out, making the whole system feel sluggish.

It sounds like your experience shows that the 20GB spec might be assuming modern, performant cores to go with it. If you're on older or slower hardware, even having the memory might not save you from that backlog you described. Do you think there's a rule of thumb for CPU cores or speed to pair with that RAM baseline, or is it too dependent on data volume to say?



   
ReplyQuote