You're right about the network, but the queue depth on the storage is the real killer. It's not just shard relocations, it's any recovery. If your storage can't drain the queue, the whole node looks unresponsive and drops out of the cluster.
Provisioned IOPS are a start, but you have to monitor `iowait` on the pods themselves. The cloud's volume metrics lie.
SQL is enough
Monitoring `iowait` is the correct diagnostic, but it's often too late. By the time you see sustained iowait, the queue is already saturated and recovery has likely begun. You need to track the pending queue length on the volume itself, alongside the ES node's indexing queue, to get ahead of it.
The cloud metrics lie because they average over a minute. A burst of indexing that fills the queue in two seconds can stall the node, but the averaged IOPS graph will look fine. We had to write a custom exporter that sampled the storage queue depth every five seconds to correlate with our Elasticsearch thread pool rejections.
That correlation revealed our actual bottleneck: the default Kubernetes `fs.inotify.max_user_watches` setting was too low, causing the node to miss file system events and leading to false storage saturation alerts. It wasn't the IOPS, it was the kernel.
Data first, decisions later.