Skip to content
Notifications
Clear all

Has anyone tried running Elastic Security on K8s at scale? Node sizing tips?

17 Posts
17 Users
0 Reactions
15 Views
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

You're right about the network, but the queue depth on the storage is the real killer. It's not just shard relocations, it's any recovery. If your storage can't drain the queue, the whole node looks unresponsive and drops out of the cluster.

Provisioned IOPS are a start, but you have to monitor `iowait` on the pods themselves. The cloud's volume metrics lie.


SQL is enough


   
ReplyQuote
(@elliotn)
Reputable Member
Joined: 3 months ago
Posts: 291
 

Monitoring `iowait` is the correct diagnostic, but it's often too late. By the time you see sustained iowait, the queue is already saturated and recovery has likely begun. You need to track the pending queue length on the volume itself, alongside the ES node's indexing queue, to get ahead of it.

The cloud metrics lie because they average over a minute. A burst of indexing that fills the queue in two seconds can stall the node, but the averaged IOPS graph will look fine. We had to write a custom exporter that sampled the storage queue depth every five seconds to correlate with our Elasticsearch thread pool rejections.

That correlation revealed our actual bottleneck: the default Kubernetes `fs.inotify.max_user_watches` setting was too low, causing the node to miss file system events and leading to false storage saturation alerts. It wasn't the IOPS, it was the kernel.


Data first, decisions later.


   
ReplyQuote
Page 2 / 2