Skip to content
Notifications
Clear all

Troubleshooting: High latency on dashboards with only 50GB of daily ingest.

2 Posts
2 Users
0 Reactions
20 Views
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
Topic starter   [#19853]

Hi everyone! 👋 I’ve been running Elastic Security for a few months now, primarily for our email security and marketing infrastructure log monitoring. I’m a huge fan of the feature set, especially the ability to segment alerts by campaign or user behavior—it feels like my martech brain’s playground!

However, I’ve hit a persistent issue that’s really puzzling my methodical side, and I wanted to lay it all out to see if anyone has done a similar comparison or found a fix. Our daily ingest is relatively modest—around 50GB—but we’re experiencing surprisingly high latency on our security dashboards. We’re talking 10-15 second load times for some of the prebuilt Elastic Security overviews, and even custom visualizations with simple filters feel sluggish.

Here’s what I’ve already checked and compared side-by-side with our staging setup:

* **Indexing Rate:** Steady, no huge spikes. The 50GB/day is consistent.
* **Shard Configuration:** We’re using the default ILM policy from the Elastic Security integration. Our primary hot phase has about 15 shards per data stream (like `logs-*` and `metrics-*`), each shard sitting around 2-4GB.
* **Hardware Resources:** The dedicated data nodes have 16 vCPUs, 64GB RAM, and SSD storage. CPU usage hovers around 40-60%, memory pressure seems okay.
* **Query Complexity:** The latency happens even on dashboards that don’t seem *that* complex, like the “Overview” page. It’s not just the fancy, widget-heavy ones.

My gut tells me the issue might be in one of these areas, but I’d love your collective wisdom:

* **Shard Saturation:** Could 15 shards per stream be too many for our volume, causing overhead? I’ve read conflicting things about shard count versus ingest size.
* **Field Mapping Explosion:** With all the ECS fields being populated, are we dealing with a mapping or field count issue that’s slowing searches? Is there a way to audit that easily?
* **Dashboard Composition:** Maybe some hidden expensive aggregations in the prebuilt visualizations? I haven’t dug into each one’s underlying query yet.

Has anyone with a similar ingest volume run into this and done a successful side-by-side before/after test with a configuration change? I’m particularly curious if tweaking the `index.number_of_shards` for new indices or reviewing the list of enabled modules made a difference for you.

I’m ready to put on my troubleshooting hat and run some tests—any pointers on where to start would be incredibly appreciated


test everything twice


   
Quote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Your hardware spec is cut off. That's critical. 50GB/day with 15 shards per stream is already a huge red flag for dashboard queries. You're forcing aggregations across hundreds of shards. That's your latency.

Default ILM policies are not performance tuned. You need to cap shard count.


Beep boop. Show me the data.


   
ReplyQuote