Hello everyone. I've been evaluating our observability and security tooling stack, and our LogRhythm agreement is up for renewal soon. While it has served certain purposes, our team is feeling the strain of its complexity and cost, especially for the breadth of features we actually use. We're primarily using it for security log aggregation and some compliance reporting, but the operational overhead is significant.
We are explicitly **not** interested in Splunk (cost and operational model) or QRadar (complexity feels similar to our current state). Our core needs are:
* **Centralized log aggregation** for security events (Windows Event Logs, firewall, IDS/IPS) and key application logs.
* **Performant search and correlation** for incident investigation.
* **Reliable alerting** with decent customization to reduce noise.
* **A more modern, lower-overhead operational model** (SaaS or self-hosted but easier to maintain).
* Clear, predictable pricing tied to our actual usage volume.
I've started a preliminary analysis and would love the community's real-world experiences. Here are the alternatives I'm compiling, but I'm keen to hear about your production runbooks and any incident postmortems that involved these tools.
**Primary Contenders:**
* **Elastic Stack (ELK)**: This is the open-core route. We'd likely need Elastic Security or a third-party SIEM layer on top. The upside is control and potential cost savings; the downside is the "build-your-own" operational burden.
* Key Question: For those running a security-focused ELK stack, how much dedicated FTE time is required to keep it effective (pipeline management, tuning, updates)?
* Example of a concern: We'd need to build our own alerting runbook. Here's a skeleton of a detection rule we'd have to develop and maintain in Elastic:
```json
{
"rule_id": "suspicious_network_service_stopped",
"risk_score": 47,
"severity": "medium",
"description": "Detects stopping of critical network services (e.g., firewall, logging) on Windows systems.",
"query": "event.category:process and event.action:"process_stopped" and process.name:("MsMpEng.exe", "svchost.exe") and process.args:("WinDefend", "mpssvc")",
"interval": "5m"
}
```
* **Graylog**: Often seen as more operationally straightforward than a full ELK deployment. Its focus is log management with alerting and dashboards.
* Key Question: How does it hold up for true security incident response, especially when you need to pivot across different data sources quickly?
* **Humio/LogScale**: Now part of CrowdStrike. Its pricing model (based on daily peak ingestion) is interesting, and the query performance is highly praised.
* Key Question: How is the learning curve for the query language compared to traditional SQL or Lucene? Any pitfalls in their billing model?
* **Microsoft Sentinel**: A natural candidate given our heavy Azure presence. The connector ecosystem is robust, and the SOAR playbooks are attractive.
* Key Question: How does cost predictability pan out? With data ingestion and analytics rules, have you experienced unexpected cost spikes during major incidents?
**Evaluation Criteria We're Using:**
* **Ingestion & Storage Cost**: Not just per GB, but how costs scale during incident surges.
* **Mean Time to Acknowledge (MTTA)**: How quickly can an on-call engineer understand the alert from this tool?
* **Operational Load**: SLO for tool maintenance itself. We don't want a tool that becomes a source of its own incidents.
* **Correlation Capability**: Can it replace the core correlation searches we've built in LogRhythm without excessive custom work?
If you've led a migration away from LogRhythm, what were the hidden pitfalls? What's your postmortem verdict a year later? I'm particularly interested in any chaos engineering you might have done – like deliberately spiking log volume – to test the resilience and cost of these platforms.
pagerduty certified lifer