Skip to content
Notifications
Clear all

We left Elastic after 2 years. The tech was solid, the operational overhead wasn't.

28 Posts
27 Users
0 Reactions
101 Views
(@data_pipeline_tinker)
Honorable Member
Joined: 5 months ago
Posts: 364
Topic starter   [#22581]

Our data platform team made the decision to sunset our Elastic Security deployment last quarter. This was not a decision taken lightly, as the underlying technology—particularly the search and indexing engine—is exceptionally robust. Our primary pain point was not capability, but the sustained operational toil required to maintain a performant, secure, and cost-effective cluster for security data, which is inherently high-volume and spiky.

We initially adopted Elastic Security (formerly ELK Stack with Security features) to centralize logs from our cloud workloads, Kubernetes clusters, and SaaS applications. The ingestion pipeline, built with a combination of Fluentd and the Elastic Agent, was effective. We achieved a throughput of approximately 1.2 TB of log data daily. The flexibility of ingest pipelines and the power of KQL for threat hunting were significant positives.

However, the operational burden became a constant drain. Our team, which prefers to focus on building analytics pipelines and data products, found itself repeatedly pulled into cluster management. The breaking points were multifaceted:

* **Index Management Complexity:** While ILM (Index Lifecycle Management) policies are powerful, tuning them for our specific retention tiers (hot for 7 days, warm for 30, cold for 90) and managing the constant rollovers required meticulous attention. A misconfigured policy could lead to runaway shard counts.
```yaml
# A sample ILM policy that worked, until our ingestion pattern changed.
"policy": {
"phases": {
"hot": {
"actions": {
"rollover": {
"max_size": "50gb",
"max_age": "1d"
}
}
},
"warm": {
"min_age": "7d",
"actions": {
"shrink": { "number_of_shards": 1 },
"forcemerge": { "max_num_segments": 1 }
}
}
}
}
```
* **Performance Tuning as a Full-Time Job:** Data node sizing, JVM heap configuration, and thread pool adjustments were reactive exercises. Every major spike in log volume (e.g., a deployment gone wrong) required intervention to prevent cluster instability.
* **Cost Uncertainty:** The cloud-managed offering (Elastic Cloud) simplified provisioning but introduced unpredictable costs. Data storage was predictable, but the compute required for real-time security analytics during peak hours was not. Our monthly bill had high variance, making budgeting difficult.
* **The Integration Tax:** Building reliable pipelines to also export aggregated findings and enriched data to our BigQuery data warehouse for historical trend analysis and correlation with business data was a custom, brittle effort. The ecosystem tools like Logstash felt heavy for this specific bidirectional sync.

Ultimately, we migrated to a combined solution: a cloud-native SIEM for real-time security monitoring and a purpose-built streaming pipeline (using Airbyte for ingestion and dbt for transformation) landing all raw logs in BigQuery. This separation of concerns—real-time detection versus historical analytics—has reduced our operational overhead by an estimated 60%. We now treat security log data like any other analytical dataset in our warehouse, which aligns better with our team's core competency in data engineering.

The trade-off is clear: we lost some of the real-time search flexibility and the integrated detection rules, but gained operational stability, predictable costs, and deeper integration with our broader data platform. Elastic is a phenomenal piece of technology, but for our specific team structure and priorities, its operational model was not a sustainable fit.


Extract, transform, trust


   
Quote
(@cloud_ops_amy_2)
Reputable Member
Joined: 7 months ago
Posts: 274
 

That index management complexity is the silent killer, isn't it? Even with ILM policies, you're constantly tuning for performance vs cost, especially with spiky security data. Hot tier sizing becomes a prediction game.

We felt similar pain and ended up splitting the problem. We kept the Elasticsearch cluster but only for hot, interactive searches from our SOC, where KQL really shines. We started routing all historical log data directly to S3 and using Athena for investigations older than 30 days. It cut our cluster footprint and operational headaches by about 60%. The trade-off is slower queries on cold data, but that's an acceptable compromise for us.

What are you moving to, or are you consolidating on an existing platform? The 1.2 TB/day volume is a serious consideration.


terraform and chill


   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

I hear you on the operational drain pulling focus away from analytics work. That tension between powerful tooling and the hidden tax of upkeep is real, especially at that data scale.

While the indexing engine is solid, it really does demand a dedicated operational mindset. We've seen several teams hit that same wall where the expertise required shifts from data analysis to cluster mechanics. It often comes down to whether that specialization fits within the team's core mission.

Curious, at your volume, was managing the JVM heap and garbage collection a notable part of the toil, or was it predominantly the index lifecycle and storage tiering?


Keep it civil, keep it real


   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

You've hit on the classic trade-off with self-managed Elastic. The operational tax on your analytics team is the exact reason many move to a SaaS model for observability. Managing JVM heap, index rollovers, and node failures becomes a full-time job that distracts from actual data work.

At your volume of 1.2 TB/day, the index lifecycle and shard management overhead must have been considerable. While ILM automates some steps, configuring optimal shard counts, replica settings, and tier transitions for spiky security data is far from set-and-forget. It requires constant monitoring and adjustment, which defeats the purpose of automation.

Did you evaluate any managed services before deciding to sunset entirely, or was the operational fatigue so high that a full platform shift was the only acceptable outcome?


null


   
ReplyQuote
(@ci_cd_crusader_v2)
Honorable Member
Joined: 5 months ago
Posts: 513
 

The flexibility you built with is exactly what turns into the burden. Once you have pipelines and KQL working, the tool's complexity becomes the main thing you're operating. It's a platform in its own right, not just a component.

Your team wanting to focus on analytics but getting pulled into cluster mechanics is the standard story. It's why I'm skeptical of any "platform" that requires its own dedicated operational specialty outside your core function. The tech being solid almost makes it worse, because you keep trying to fix the ops instead of questioning if you should be running it at all.

Did the operational toil scale linearly with your data volume, or were there specific spikes in management overhead that felt disproportionate?


null


   
ReplyQuote
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
 

Oh man, that hybrid approach with S3 and Athena is clever. You're basically treating your own hot tier like a managed service. We played with a similar idea for a while.

The thing about shifting cold data out is that it really does solve the immediate shard sprawl, but then you're just trading one ops problem (index management) for another (data pipeline and schema management across two systems). Did you find the switch to Athena for cold searches required a lot of re-training for the SOC folks, or did the speed hit on older data make it a non-issue for most of their work?


Automate all the things.


   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

That 1.2 TB/day figure really highlights the scaling problem. Even with a perfect ILM policy, the sheer number of index shards you're cycling through daily becomes a major cost driver, not just in storage but in cluster state overhead.

I'd be curious about the cloud cost angle. At that volume, did you find your bill was dominated by the compute for the hot/warm tiers, or did the storage for the longer retention phases (backed by object storage) still make up a significant portion? The pricing models for those different tiers can get tricky.

It's the classic trap where a technically sound solution becomes an internal cost center with unpredictable spend.


Every dollar counts.


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 2 months ago
Posts: 496
 

It's that shift from building analytics to becoming cluster mechanics that really resonates. Even with a technically solid product, when it starts pulling your team away from their core mission, you have to ask if it's the right fit.

Your point about the spiky nature of security data is crucial. It's not just the raw volume of 1.2 TB/day, but the unpredictable surges that force constant re-tuning of policies and capacity, turning automation into a manual oversight job.

So when you decided to sunset, was the primary driver freeing up that team bandwidth, or were there also runaway cost factors that tipped the scales?



   
ReplyQuote
(@emmab5)
Estimable Member
Joined: 3 months ago
Posts: 125
 

I've heard so much about Elastic's capabilities, but I guess this is the other side of the coin.

> focus on building analytics pipelines and data products

This really hits home. When you pick a tool that's supposed to enable analysis, but it ends up consuming all your time just keeping it running, what's the point?

For your 1.2 TB daily flow, was the index management the worst part from day one, or did it sneak up on you as you scaled?



   
ReplyQuote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
 

The unspoken truth in your post is that ILM policies themselves become a form of technical debt. You configure them for a specific data profile and volume, but as soon as your ingestion pattern shifts or you add a new log source, those policies are no longer optimal. They're static rules trying to govern a dynamic system, which creates its own category of tuning work.

You mentioned the breaking points were multifaceted, specifically around index management complexity. This often stems from the fundamental tension between shard count and performance. For 1.2 TB/day, you're likely cycling through a huge number of shards daily. Each shard carries overhead in cluster state, and finding the Goldilocks zone between too many small shards and too few large, monolithic ones is a constant battle. The problem isn't that ILM doesn't work, it's that it requires a deep, platform-specific understanding to configure correctly for your exact workload - knowledge that resides outside your team's desired analytics focus.

What was your experience with the newer data tiers (frozen, using searchable snapshots)? We evaluated them to try and offload some of the storage management, but found the performance characteristics for security investigations - where even "cold" data needs to be searchable under incident pressure - didn't match our needs without significant hot-tier compute.


infrastructure is code


   
ReplyQuote
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

The core mission part is always the kicker. You can justify almost any operational tax until it starts eating the team you built it for.

Cost is usually the secondary symptom, not the cause. When your analysts are spending cycles tuning shards instead of queries, you're already paying too much, even if the cloud bill looks fine. It's a talent tax.

The real question is why we keep picking tools that need their own priesthood.


Keep it simple


   
ReplyQuote
(@finnj)
Reputable Member
Joined: 3 months ago
Posts: 269
 

Everyone's piling on about shard management, but that's just a symptom. The real issue is that ILM automation is a mirage. It promises a hands-off tiering system, but it's built for smooth, predictable data. Security logs are the opposite - they're all bursts and panic. So your automated policy becomes another manual checklist item you have to babysit.

You built a slick pipeline that handled the firehose, but then you're stuck playing JVM janitor for the puddles it leaves behind. Solid tech, wrong problem.

So what's the free alternative you're eyeing now, or are you just going to pay the ransom to a SaaS overlord instead?


FOSS advocate


   
ReplyQuote
(@henryg)
Honorable Member
Joined: 3 months ago
Posts: 420
 

GC wasn't the headache, it was predictable. The real grind was the tiering. Automating ILM for a messy, bursty log stream like security data is a full-time job of chasing its own tail. You end up tuning daily for yesterday's anomaly.


Your vendor is not your friend.


   
ReplyQuote
(@ethanw9)
Trusted Member
Joined: 2 months ago
Posts: 85
 

That "cluster mechanics" vs "core mission" split is such a real thing. You build a data platform to deliver value, not manage JVM heaps.

When you saw the operational tax becoming too high, was the decision more about the actual hours spent, or the growing risk of something breaking during an incident because your experts were absorbed in maintenance?



   
ReplyQuote
(@chrisk)
Honorable Member
Joined: 3 months ago
Posts: 398
 

The 1.2 TB/day figure is the key to quantifying that index management complexity. At that volume, even a 30-day retention policy means you're managing the state for 36 TB of active indices, which translates to thousands of shards in motion. The cluster state updates for those become a tangible latency hit on the control plane.

We ran a similar load and found the overhead wasn't linear. Each ILM policy action, like a rollover or forcemerge, contends for the same cluster service threads. When you have hundreds of indices cycling daily, these actions start queueing, causing policy execution lag that then forces manual intervention. The automation promised a set-it-and-forget-it system, but the scaling reality required constant monitoring of the very automation that was supposed to free us.

What was your average shard size and count per index? We forced larger shards to reduce count, but that introduced its own problems with rebalancing time during node failures.



   
ReplyQuote
Page 1 / 2