Skip to content
Notifications
Clear all

QRadar deployment gone wrong - what we learned the hard way

2 Posts
2 Users
0 Reactions
3 Views
(@new_evaluator_2025)
Eminent Member
Joined: 4 months ago
Posts: 16
Topic starter   [#2017]

Hi everyone. I’ve been reading through the reviews here for weeks while we were evaluating SIEM options, and honestly, I’m still a bit overwhelmed. We ended up going with QRadar, and… well, our deployment didn’t exactly go smoothly. I wanted to share what happened, partly to vent and partly to see if anyone else ran into this stuff.

Our main goal was getting better visibility into our cloud workloads and some on-prem legacy systems. The sales process was fine, but the moment we got into the actual deployment, we hit a wall that wasn't really covered in the demos. Our assumption was that the virtual appliance would be relatively straightforward to scale. The biggest "hard way" lesson was around resource estimation. We based our initial VM sizing on IBM's datasheets, but the actual event per second (EPS) load from our log sources was way spikier than we modeled. The system choked during our peak business hours, and we spent days just trying to get the buffers under control. We learned that "capacity planning" for QRadar isn't a one-time thing; you really need to monitor its own performance metrics from day one.

Also, the pricing model bit us a bit. We thought we had a handle on the cost based on the EPS tier we licensed, but some of the cloud sources we added seemed to generate more "events" than we anticipated, pushing us close to the next pricing bracket. It felt like we were being penalized for successfully onboarding more log sources, which was the whole point!

I'm curious if others have had similar experiences. Specifically:
* Did you find the initial resource guidance from IBM accurate, or did you have to over-provision?
* For those using it in hybrid environments, were any specific log sources (like AWS CloudTrail or Azure logs) unexpectedly heavy on EPS?
* How do you handle the cost creep? Is it just about constantly tuning and filtering logs before they even hit QRadar?

We're through the worst of it now, and it's running, but the journey was way more stressful and expensive than we budgeted for. Hoping our mistakes can help someone else.


Help me decide


   
Quote
(@sre_shift_worker)
Eminent Member
Joined: 3 months ago
Posts: 23
 

Oh man, the datasheet to reality gap is a classic. The EPS spikes are what kills you. You can't just average it out, you have to plan for the 95th percentile at a minimum, otherwise your queues fill and it's game over.

We put a tiny Telegraf agent on the QRadar VM itself to pull its own system metrics into our monitoring stack. CPU wait on the disk, memory free, load average. You have to watch it like any other critical service, because by the time QRadar's own health alerts fire, you're already on the back foot.

And yeah, wait for the licensing surprise when you suddenly need more EPS capacity than you bought. Good times.


Pager duty is not a hobby


   
ReplyQuote