We just finished a 12-month evaluation of both platforms across two separate 100-node environments (mix of cloud VMs and bare metal). The goal was full SIEM/XDR coverage with a hard cap on operational overhead. Here's the raw breakdown.
**Elastic Security (Elastic Stack + paid Enterprise subscription)**
- **Deployment:** Managed with Ansible and Helm. The Elastic Agent is clean, but the Elasticsearch cluster is the resource hog.
- **Licensing Cost:** ~$45k/year list for the Enterprise features we needed (threat detection, ML). Could be lower with negotiation.
- **Infra Cost (Cloud):** ~$2.8k/month for a 3-hot/2-warm/2-ml node cluster (i3.2xlarge). This is non-negotiable for performant indexing at this scale.
- **Admin Overhead:** 15-20 hours/week for tuning detections, managing index lifecycle, and keeping the cluster healthy. The learning curve is steep.
- **Biggest Win:** Integrated searches across logs and security events are unbeatable. KQL is powerful.
- **Biggest Pain:** You are now running a critical, large-scale search database. Performance tuning is a constant battle.
**Wazuh (Open Source, self-managed)**
- **Deployment:** Single all-in-one manager or scaled with indexer/manager/dashboard nodes. Used their OVA for POC, then Terraform/Ansible for prod.
- **Licensing Cost:** $0 for core features. Paid support optional (~$15k/year).
- **Infra Cost (Cloud):** ~$800/month for 3 manager/indexer nodes (m5.2xlarge).
- **Admin Overhead:** 5-10 hours/week. Most time spent on rule tuning and integrating new log sources. The stack is simpler by design.
- **Biggest Win:** It's a dedicated security tool. The HIDS and PCI-DSS/SOC2 compliance mapping out-of-the-box saved weeks of work.
- **Biggest Pain:** The interface and alerting feel dated. Scaling the indexer past this node count requires careful planning.
**TCO Verdict for a 100-server environment:**
If you have dedicated Elasticsearch expertise in-house and need deep, ad-hoc investigation capabilities, Elastic can justify its cost. For a team that needs a focused security monitoring and compliance tool with predictable overhead, Wazuh wins on pure TCO.
Our team chose Wazuh. The deciding factor was operational simplicity. We couldn't justify the additional ~$35k/year in cloud costs plus 15+ hours/week of senior DevOps time just to keep Elasticsearch alive. Wazuh does 80% of what we needed for 30% of the total cost.
-shift
shift left or go home
I lead platform SRE for a 450-node fintech environment running a mix of on-premise Kubernetes and cloud VMs. We evaluated both solutions extensively two years ago and currently run Wazuh in production, having shifted from a paid Elastic Security deployment after a 9-month parallel run.
1. **Total Operational Overhead**
Elastic requires a dedicated platform team. You quoted 15-20 hours weekly for upkeep; that aligns with our experience. The Elasticsearch cluster itself becomes a primary service requiring monitoring, scaling, and disaster recovery. Wazuh's overhead is approximately 8 hours weekly, mostly for rule updates and alert tuning, because you're managing an appliance, not a distributed database.
2. **Infrastructure Cost Breakdown**
Your $2.8k/month cloud estimate for Elastic is accurate for warm data retention. The hidden cost is the operational burden of the underlying compute, which at 100 servers translates to 30-40% of your total node resources just for the observability stack. Wazuh's manager, indexer, and dashboard components can run on three modest 4-core/16GB VMs. Our on-premise hardware cost for those was under $15k capex, with negligible ongoing cloud spend.
3. **Detection Tuning and Rule Management**
Elastic's detection rules are powerful but exist in a proprietary DSL. Modifying or creating complex correlation rules requires deep platform knowledge. Wazuh uses a decoupled, XML-based rule system. While less modern, it allows direct git-based versioning and merging of community rules from its public repository. We sync 300+ custom rules from a GitOps pipeline with no service restart.
4. **Scalability and Performance Ceiling**
Elastic scales horizontally by design and can handle the 100-server load easily, but you must provision for peak ingestion, not average. Wazuh's indexer component can become a bottleneck if event volume exceeds 5,000 EPS consistently per manager, requiring a scaled deployment. At 100 servers, assuming normal log levels, a single indexer node is sufficient but requires monitoring of the queue depth.
Given your hard cap on operational overhead and the 100-server scale, I'd recommend Wazuh for a team that prioritizes operational simplicity and predictable costs. The trade-off is accepting a less integrated search experience compared to KQL. If your team has strong Elasticsearch operational expertise and the budget for dedicated platform management, Elastic Security provides a superior investigative surface. To make a clean call, specify if you require real-time compliance reporting (like PCI continuous monitoring) and the average EPS you expect from your busiest 10 servers.
Data over dogma