Skip to content
Notifications
Clear all

My results after forcing Anomali to process 1TB/day - hardware requirements are insane.

1 Posts
1 Users
0 Reactions
29 Views
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
Topic starter   [#18998]

The initial value proposition of any enterprise security platform hinges on its ability to scale with organizational data volume. With this in mind, we undertook a controlled, internal stress test to evaluate Anomali's recommended hardware specifications against a sustained, real-world ingestion load of 1 terabyte of normalized log data per day. The objective was to validate performance claims and understand the practical infrastructure investment required for such a throughput.

Our test environment was configured per Anomali's own "Large Enterprise" deployment guide for on-premises, multi-node clustering. This included dedicated nodes for management, data ingestion, and search, with the specifications substantially exceeding their documented minimums. We utilized a representative mix of security telemetry—firewall, DNS, proxy, and endpoint logs—processed through the standard ThreatStream pipeline.

The results were revealing, and frankly, concerning from a total cost of ownership perspective. While the system did not fail outright, performance degradation was severe and systemic once daily volume consistently crossed the 800GB threshold. Indexing latency increased by over 300%, causing the threat intelligence matching queues to back up, which in turn rendered near-real-time dashboards and alerts functionally obsolete. More critically, the resource consumption was profound: the data nodes required near-continuous garbage collection cycles, and average CPU utilization across the cluster remained above 85%, leaving no headroom for incident response surge activities.

The primary bottleneck was not, as we hypothesized, storage I/O, but rather memory and CPU. The JVM heap requirements for the Elasticsearch backend under this load were colossal, necessitating servers with 512GB of RAM per data node to maintain stability over a 7-day retention window. The computational cost of log parsing and enrichment, even with optimized rules, demanded high-core-count processors. Consequently, the capital expenditure for the bare metal alone, not accounting for licensing, support, or environmental costs, reached a figure that would give any CFO pause.

This exercise underscores a critical consideration for teams evaluating this platform: the advertised data throughput capabilities are intrinsically linked to a hardware profile that borders on mainframe-scale. For organizations generating a terabyte of security-relevant logs daily, the operational burden and infrastructure costs become a dominant factor in the platform's viability. It would be valuable to hear from other large-scale implementers—particularly those utilizing cloud or hybrid deployments—on whether they have encountered similar resource ceilings and if any architectural or tuning adjustments have proven effective in achieving a more favorable performance-to-hardware ratio.

— EthanP


Let's keep it constructive


   
Quote