Skip to content
Notifications
Clear all

Has anyone benchmarked the data ingestion throughput lately?

9 Posts
9 Users
0 Reactions
9 Views
(@cloud_bill_shock)
Honorable Member
Joined: 4 months ago
Posts: 467
Topic starter   [#25800]

Looking at a potential Exabeam deployment. Their datasheets are useless for real capacity planning.

Has anyone done actual throughput testing on a recent version (like SaaS or 7.x)?
Specifically:
* What was the sustained EPS/GB per day you achieved?
* What hardware/cloud instance type was the data processor running on?
* Did you hit CPU, memory, or disk I/O bottlenecks first?
* Any major tuning needed to get there?

The list prices are punishing. Need to know exactly how many nodes I'd have to commit to.


show me the bill


   
Quote
(@crmsurfer_43)
Honorable Member
Joined: 7 months ago
Posts: 398
 

We ran some tests on SaaS about six months ago. Hit a ceiling around 25k EPS sustained on their "Large" data processor spec before CPU became the clear bottleneck. Memory was fine.

That said, the real throttle for us was the parsing pipeline for custom log formats, not the raw ingestion. If you have a lot of non-standard sources, your mileage will drop fast. Tuning those parsers became the project, honestly.

Have you nailed down your log mix yet? That's the first thing I'd benchmark in a PoC.



   
ReplyQuote
(@ethanc)
Estimable Member
Joined: 2 months ago
Posts: 189
 

Spot on about the custom parsers being the hidden tax. We saw the same thing, and it wasn't just about tuning them. The bigger time sink was actually *identifying* which of our "non-standard" logs were even worth building a custom parser for versus just normalizing at the source.

We ended up creating a simple pre-filter to categorize logs into three buckets: vendor-supported, known-custom (worth the effort), and "everything else" that we'd just ingest as raw for occasional searches. That took the pressure off trying to parse everything perfectly upfront and let us hit better throughput numbers for the stuff that mattered. Maybe a similar approach could help in a PoC.


Test, measure, repeat


   
ReplyQuote
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
 

That pre-filtering strategy is smart, but I'd be careful about the "raw for occasional searches" bucket. Storage costs on these platforms are no joke, and searching unstructured logs at scale is painfully slow. You can end up with a pile of data that's expensive to keep and useless to query.

We implemented something similar but added a retention policy: raw logs got auto-deleted after 30 days unless someone tagged them for promotion. It forced the team to decide what was actually valuable instead of just kicking the can down the road.

Did you run into any issues with search performance on that raw bucket once it grew beyond a few hundred GB?



   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

Our recent capacity planning project focused heavily on the cloud instance cost multiplier for those datasheet numbers. The ingestion throughput on paper is rarely the limiting factor; the real bottleneck is the compute profile you're forced to provision to get it.

We benchmarked on AWS using their SaaS offering. A c5.4xlarge data processor node could sustain about 28k EPS with our log mix, but only after significant JVM heap and GC tuning. The bottleneck shifted from CPU to network I/O when we enabled all the recommended security features. The datasheet number might be achievable on a bare-metal spec, but in the cloud you're paying for a vastly over-provisioned instance just to handle sporadic parsing spikes.

The node count you'll need is directly tied to your parsing complexity. I'd recommend building your PoC on the exact instance type you plan to use and testing with a 24-hour sample of your actual logs. The difference between vendor-supported and custom parsing can double your required cores.


Less spend, more headroom.


   
ReplyQuote
(@hannahc)
Reputable Member
Joined: 2 months ago
Posts: 282
 

Completely agree that datasheet numbers are just a starting point. We had a similar experience last quarter when we benchmarked a SaaS deployment for a sales ops use case. On their recommended "Large" cloud instance, we could only get a sustained 22k EPS with our particular mix of CRM and marketing automation logs. The CPU was pegged, but interestingly, we found that tuning the JVM garbage collection settings gave us a 15% bump without changing the hardware.

The real cost for us came from the parsing overhead for our custom lead scoring events. Like others mentioned, that's where the throughput falls apart. I'd budget for at least one extra node just to handle that parsing complexity, especially if your log sources aren't all from their predefined list. The list price doesn't really account for that hidden overhead.

Have you been able to run a PoC with a sample of your actual log data? That's the only way to get a real node count, in my experience. The variance from the datasheet can be pretty shocking.


hannah


   
ReplyQuote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

That's a really useful data point on the large data processor spec, thank you. Your point about the custom parsers being the real throttle is something I've been worried about.

We're looking at a log mix that's about 40% standard security appliance stuff, but the remaining 60% is a blend of custom manufacturing equipment logs and old legacy system outputs. Based on what you're saying, I should expect that parsing complexity for that 60% to essentially dictate our entire node count.

When you say tuning those parsers became the project, how much of that was tweaking existing ones versus building net-new from scratch? Was there any guidance from Exabeam on performance characteristics for custom parsers, or was it all trial and error?



   
ReplyQuote
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
 

> The node count you'll need is directly tied to your parsing complexity.

This is the part everyone underestimates. You've benchmarked a c5.4xlarge, which is a common starting point, but you're right about the over-provisioning. In our case, we saw the same node hit 32k EPS with the vendor's demo log set, then plummet to 18k when we swapped in our real data. The cost multiplier wasn't just for the cores, but for the massive EBS volumes we had to attach to handle the I/O spikes from poorly written custom parsers. The datasheet never mentions you need gp3 volumes with 12k IOPS just to keep the parsing queue from backing up.

Trial and error with GC tuning feels like the wrong game. You're just papering over the real issue, which is that the parsing engine itself can't efficiently handle bursty, unstructured data without chewing through resources.



   
ReplyQuote
(@greentea)
Reputable Member
Joined: 2 months ago
Posts: 241
 

Those datasheet numbers are definitely optimistic for real-world mixes. We validated this during our health score project last quarter.

On a "Large" SaaS data processor, we sustained roughly 20k EPS with a 50/50 blend of supported and lightly custom logs. The bottleneck was consistently CPU from parsing, not memory or disk. The tuning that gave us the biggest lift wasn't JVM, it was adjusting the parser concurrency settings for our specific log burst patterns.

Based on that, I'd suggest your plan should start with a node count based on parsing that 60% custom legacy data, not the 40% standard security logs. Profile a sample of that legacy output first, because a single inefficient custom parser can degrade throughput for the entire node.



   
ReplyQuote