Skip to content
Notifications
Clear all

Results after 6 months: ES cost us 2.5 FTE to maintain, was it worth it?

19 Posts
19 Users
0 Reactions
30 Views
(@andrewb)
Reputable Member
Joined: 3 months ago
Posts: 292
 

You're assuming manual hunts are the only alternative. That's exactly how these vendors frame the problem to trap you. There are plenty of middle-ground tools that automate detection without needing a team of log janitors.

> operational overhead from infrastructure drift

The tool is supposed to be the solution to this, not the source of it. If a minor OS patch breaks your aggregations, your detection logic is just a house of cards built on vendor magic. Maybe that 2.5 FTE is the real cost of pretending parsed fields are trustworthy.


—aB


   
ReplyQuote
(@devops_shift_lead)
Honorable Member
Joined: 6 months ago
Posts: 443
 

> The lock-in risk also shifts from a vendor's proprietary framework to more open formats and skills.

That's the theory. In practice, you often end up locked into your own homegrown framework instead, which can be worse. Now you're the one writing the content packs and maintaining the parser treadmill. The institutional muscle is great if you can afford the gym membership. With 2.5 FTE already burned, I doubt they can.


shift left or go home


   
ReplyQuote
(@emilyr22)
Reputable Member
Joined: 3 months ago
Posts: 229
 

I've seen this play out with data pipelines for dashboards too. The central team becomes the bottleneck because they're the only ones who can fix broken transforms.

But I'm curious, if a simpler tool makes the pain obvious, how do you stop app teams from just ignoring the broken data? Doesn't that just create a different kind of alert fatigue?



   
ReplyQuote
(@gabrielm)
Reputable Member
Joined: 3 months ago
Posts: 253
 

You've put it really well. The shift from vendor-specific skills to general platform engineering is a real benefit, but I'm still unsure about the overall cost.

If you're paying platform engineers to maintain your Prometheus stack, doesn't that still require a centralized team with deep expertise? You've just traded Elasticsearch certified engineers for Kubernetes and observability specialists. Is that pool of talent actually easier to find or retain?

This makes me wonder about a direct comparison. For a team already running a complex k8s setup, is the tax for Prometheus lower than the one for ES? Or does it just feel more "modern" while consuming similar cycles?



   
ReplyQuote
Page 2 / 2