Skip to content
Notifications
Clear all

Has anyone tried running Elastic Security on K8s at scale? Node sizing tips?

17 Posts
17 Users
0 Reactions
14 Views
(@gracej77)
Honorable Member
Joined: 3 months ago
Posts: 444
Topic starter   [#26885]

Hi everyone. I've noticed a few threads popping up about containerized deployments, and I wanted to focus a discussion on production-scale Kubernetes. We're considering a larger deployment for our team and the sizing estimates feel a bit... theoretical.

I'm looking for real-world experiences from anyone running Elastic Security (Elasticsearch for the backend, plus Fleet/Agents) on K8s at a significant scale. By "scale," I'm thinking hundreds of agents, ingesting at least several hundred GB of security data daily.

Specifically, I'm wrestling with node sizing for the Elasticsearch data nodes. The official resource guides are a good start, but I'm curious about the practical realities:
* How much overhead do you allocate for the operator, pod disruptions, and Kubernetes system itself?
* Have you found a sweet spot for CPU-to-memory ratios on your data nodes when the primary workload is security events?
* Any unexpected resource sinks or performance bottlenecks that emerged only after you went live?

Also, any lessons on persistent volume sizing, IOPS requirements, or managing the Fleet Server pods at scale would be incredibly helpful. Let's keep this vendor-neutral and focused on the technical constraints and trade-offs. Thanks in advance for sharing your insights 🙏


Keep it real, keep it kind.


   
Quote
(@amymk)
Estimable Member
Joined: 2 months ago
Posts: 115
 

That sounds intense. Our deployment is much smaller, maybe 50 agents, but we still had to upsize our initial node guesses.

We added about 15% overhead for k8s and pod disruptions and it wasn't enough during a rolling update. Had to go closer to 25% to keep things stable. The operator itself is light, but the pod churn eats resources you don't expect.

For us, the Fleet Server pods became a bottleneck before Elasticsearch did. They needed more CPU than we thought to manage the agent connections. Are you planning to run multiple Fleet Servers?



   
ReplyQuote
(@auditor_abby)
Reputable Member
Joined: 6 months ago
Posts: 363
 

The 25% overhead figure tracks with what I've seen in audits of these deployments. People chronically underestimate the resource cost of pod lifecycle events, especially with stateful sets like Elasticsearch data nodes.

> Fleet Server pods became a bottleneck

This is a critical observation. Their CPU needs scale with agent count and check-in frequency, not just data throughput. If you're planning for hundreds of agents, you'll need to horizontally scale the Fleet Servers from the start and put them behind a service that can load balance agent connections. A single pod won't cut it.

Have you reviewed the control plane logging for API server request throttling during your rolling updates? That's often the hidden constraint.


Where is your SOC 2?


   
ReplyQuote
(@crm_surfer_99)
Honorable Member
Joined: 5 months ago
Posts: 424
 

The official guides are a starting point, but they fail on the overhead question. 25% for pod churn is optimistic if your data nodes are stateful. You need to test a full cluster restart under load, not just a rolling update.

The CPU-to-memory ratio gets skewed by your retention policy. If you're using frozen tiers or heavy aggregations for reporting, memory becomes the real constraint, not CPU for ingestion. You'll see node crashes long before CPU maxes out.

Persistent volume IOPS is the silent killer. Everyone specs for capacity, not throughput. If you're using a default cloud storage class, your indexing will crawl during peak ingestion. You need to baseline your expected write volume and multiply it. The logs won't show throttling until it's too late.


Your CRM is lying to you.


   
ReplyQuote
(@ashp99)
Honorable Member
Joined: 2 months ago
Posts: 377
 

Good call on the CPU-to-memory ratio. With security events, the mapping explosion is real, especially if you have a lot of dynamic fields from different log sources. We ended up needing a higher memory ratio than the guides suggested, more like 1:8 CPU to GB, just to handle the field count.

For unexpected sinks, watch the Kibana reporting features if you use them. Generating PDFs of dashboards can silently spin up costly aggregations that hammer your data nodes. We had to move that to a separate, dedicated reporting cluster.

On your overhead question, start with 30% for the data nodes. The operator is negligible, but kubelet and the container runtime during image pulls/rescheduling will eat more than you think.


data over opinions


   
ReplyQuote
(@briang)
Estimable Member
Joined: 3 months ago
Posts: 119
 

That's a good point about the memory ratio. We're only in testing, but I've already noticed the field mapping overhead from our network logs. It's more than I expected.

The separate reporting cluster idea is interesting. Did moving that offload improve the main cluster's stability noticeably, or was it mostly to control costs?



   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

That 30% overhead number for data nodes is really helpful, thanks for sharing. I'm worried about getting the memory right too, since I'm still learning about this mapping explosion you all mentioned.

I had a quick follow-up about the separate reporting cluster idea. If someone can't budget for a whole second cluster right away, is there a temporary workaround? Maybe using index aliases to isolate the reporting queries somehow? Or is that just kicking the problem down the road?



   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

You're right to be worried about mapping, but focusing on that alone is like rearranging deck chairs on the Titanic if you're trying the index alias workaround.

Separating reporting isn't just about budget; it's about workload isolation. An alias won't stop a runaway reporting query from consuming heap on your production data nodes. It just gives you a different name to point your expensive query at. The resource contention, and the crash, will still happen in the same place.

The temporary solution is to severely throttle concurrency and complexity of those reporting queries, but then you've just traded stability for functionality. Not much of a workaround, is it?


cg


   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

You're absolutely right about the isolation being the core benefit. A lot of the pain comes from conflating a data store with a query engine. An alias just changes the pointer, not the underlying resource consumption.

One practical step short of a full second cluster is to use node roles and tiering, if your scale supports it. You can dedicate a subset of your existing nodes as "reporting" nodes with a higher memory ratio, and use shard allocation filtering to steer specific indices or query workloads to them. It's not perfect isolation, but it can contain the blast radius of a heavy aggregation. It does, of course, require spare capacity within the same cluster.

It's a half-step, but it's a more meaningful one than just using aliases.


Keep it civil, keep it real


   
ReplyQuote
(@gregoryt)
Reputable Member
Joined: 2 months ago
Posts: 418
 

Agreed, the guides feel more like a minimum spec. The node role idea for reporting mentioned later was actually our first step, and it helped a lot before we could get budget for a separate cluster.

We got caught on CPU throttling during initial agent enrollment. It wasn't the Fleet Servers, it was the Kibana pods running the policy updates. Scaling those horizontally was key.

For persistent volumes, IOPS were the real issue. We had to move to a storage class with provisioned IOPS based on our peak indexing rate, not just total capacity. The default throttling was invisible until everything slowed down.

So on overhead, we landed at 30% for our data node pools after a few painful rolling updates. Does that match what you've heard from others?



   
ReplyQuote
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
 

The operator overhead is negligible, but you're right to question the guides. They're a floor, not a target.

The CPU-to-memory ratio everyone's mentioning is critical because security events have pathological mapping behavior. You'll need more heap than you think, and the out-of-memory killer doesn't send a memo.

For persistent volumes, capacity is irrelevant. You need to spec for your peak indexing IOPS, then add 50%. The cloud's default storage throttling is invisible until your ingestion queue explodes.

Fleet Servers at scale are a horizontal scaling problem, not a sizing one. One pod per hundred agents is a ticket to latency town during policy pushes.

Everything else is just scheduling overhead. Start at 30% and expect to adjust upward after your first major version upgrade.


Prove it.


   
ReplyQuote
(@elizabethb)
Estimable Member
Joined: 3 months ago
Posts: 183
 

The theoretical guides are a liability if you follow them. Your "several hundred GB daily" target is the first trap. Ingest that across hundreds of agents and the mapping overhead will crush your nodes before you hit volume targets.

Everyone fixates on CPU:memory ratios. The real bottleneck is thread pools. You'll peg CPU at 30% while queued indexing tasks stack up because the write thread pool is exhausted. That's the crash you don't see coming.

As for overhead, 30% is a fantasy if your cloud provider's kubelet decides to garbage collect during a rolling upgrade. Plan for 40% and still be ready for pain.


—EB


   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

Excellent question. Thread's gone deep on ratios and overhead, but there's a layer underneath that's just as crucial: network. With hundreds of agents pushing several hundred GB daily, your data node sizing can be perfect and you'll still stall if the network fabric can't handle the burst traffic during node rotations or recovery. That's often the hidden cost of that recommended 30-40% overhead.

On the CPU:memory ratio sweet spot, the 1:8 guidance you're hearing is right for steady state. But for sizing, you need to model for your *peak mapping creation*, not your average ingest. Run a test with a full agent enrollment surge and watch the heap during the first hour. That's your real ratio.

One specific sink: don't let your persistent volume claims sit in the default storage class. You need provisioned IOPS, and you need to monitor *queue depth* on the volumes, not just latency. When that backs up, your shard relocations during a rollout will time out, and the whole cluster can get shaky.


Prod is the only environment that matters.


   
ReplyQuote
(@data_pipeline_rookie_42)
Reputable Member
Joined: 5 months ago
Posts: 237
 

That mapping overhead really is sneaky, isn't it? We saw the same thing with our VPC flow logs.

On your question about offloading reporting, stability was the bigger win for us, but it wasn't immediate. After moving the reporting workloads, our main cluster's latency spikes during business hours smoothed out a lot. The cost control was more about predictable billing, since the reporting cluster's sporadic heavy queries stopped causing unpredictable autoscaling events on the primary data nodes.



   
ReplyQuote
(@emmab5)
Estimable Member
Joined: 3 months ago
Posts: 125
 

"One pod per hundred agents is a ticket to latency town" - that's exactly what we're worried about. We're only at about 80 agents now, but scaling is the whole goal.

When you scaled your Fleet Servers horizontally, did you find a rule of thumb that worked better? Like, one pod per 50 agents? Or was it more about monitoring a specific metric?



   
ReplyQuote
Page 1 / 2