I've been tasked with evaluating our SIEM strategy as we complete a full migration to AWS. Our legacy on-premises QRadar deployment is no longer fit for purpose, but we have significant institutional knowledge with the platform. The logical question is whether to transition to a native cloud SIEM or attempt to run QRadar in a cloud-only context.
I'm looking for detailed, operational experiences from teams running QRadar entirely within a public cloud (AWS, Azure, or GCP). My primary concerns are architectural and financial:
* **Data Ingestion & Cost:** How do you manage the flow and cost of VPC flow logs, CloudTrail, GuardDuty findings, and container logs into the SIEM? Are you using native services (e.g., Kinesis, EventBridge) or the QRadar cloud appliances? What is the data transfer/egress cost profile?
* **Deployment Model:** Are you using the SaaS offering (QRadar on Cloud) or managing your own VMs with the licensed software? If the latter, how do you handle scalability and high availability compared to the on-prem model?
* **Resource Management:** For self-managed deployments, how do you right-size EC2 instances for Event Processors and Console/Ariel components? Have you explored Reserved Instance strategies for these long-running nodes to optimize commitment?
The marketing materials suggest it's possible, but I'm interested in the practical pitfalls. For instance, does the licensing model based on EPS/Flow per second become problematic or more expensive when ingesting verbose cloud service logs compared to traditional network data? Any insights on managing the appliance lifecycle (patching, upgrades) without a physical footprint would also be valuable.
Your bill is too high.
We went through this evaluation about 18 months ago and ultimately chose a native cloud SIEM, but our deep dive into QRadar cloud-only was extensive. On your first point about **data ingestion and cost**, the architecture we modeled used Kinesis Data Firehose for CloudTrail and VPC flow logs, feeding into the QRadar cloud appliance. The cost wasn't trivial. While it's more efficient than raw egress, you're still paying for Firehose processing and Kinesis shards, plus the egress from Firehose to QRadar over a private link. For container logs, we found we'd need a separate forwarder (like Fluentd) aggregating to S3, then another process to pull into QRadar. The data pipeline complexity felt like rebuilding an on-prem collection infrastructure but with cloud billing traps.
Regarding the **deployment model**, we looked hard at self-managed on EC2 because of perceived cost control. The scaling got messy quickly. In an on-prem setup, you throw hardware at it. In AWS, auto-scaling Event Processors based on EPS is theoretically possible, but the QRadar licensing tied to EPS and VM cores created a nightmare for forecasting. High availability required multi-AZ deployments and managing the QRadar HA configuration across them, which felt fragile compared to the SaaS version's managed SLAs. The operational overhead of patching, securing, and monitoring the underlying VMs brought back all the pains we were trying to leave behind.
If you have the institutional knowledge, the SaaS offering (QRadar on Cloud) might be the compromise, but you lose a lot of control over the underlying infrastructure and the cost model shifts to a fully managed service. Have you calculated the total cost of ownership comparison, factoring in the FTEs needed to manage a self-hosted cloud deployment versus the SaaS premium? That was the deciding factor for us.
Support is a product, not a department.
Ran a fully self-managed QRadar cluster on AWS for three years before switching. On your resource management question, right-sizing is a constant battle.
The Ariel EC2 sizing advice from IBM is optimistic. You'll hit I/O bottlenecks on r5 instances during complex searches long before CPU maxes out. We ended up on i3 instances with NVMe for the Ariel nodes, which helped but wrecked the cost model.
For Event Processors, the cloud environment makes auto-scaling nearly impossible due to licensing. You're stuck manually forecasting peaks and over-provisioning. A major CloudTrail spike from a misconfigured trail can stall the entire pipeline.
Five nines? Prove it.
We stuck with self-managed QRadar VMs in Azure for about two years for that exact reason, keeping our team's expertise. The data transfer costs for logs were a constant budget surprise, honestly.
You can manage it with Event Hubs and storage accounts, but the pipeline complexity adds operational overhead that eats into the time saved from not training on a new platform. For right-sizing, we had to schedule major searches for off-peak hours to avoid the I/O bottlenecks others mentioned.
Trust the trial period.
Your institutional knowledge is a sunk cost. Clinging to QRadar because your team knows it is how you end up rebuilding an entire on-prem log shoveling operation in the cloud, just with pricier, more complex plumbing.
You're asking about right-sizing EC2 instances for Ariel. That's the wrong question. The real problem is that the architecture itself is a poor fit. You'll constantly be fighting I/O bottlenecks on managed disk, leading you to specialized, expensive instances like i3. Then you're locked into that footprint because the licensing model makes actual elasticity a fantasy. You scale by manual forecast and over-provisioning, which negates a core cloud benefit.
The operational overhead of managing those data pipelines and instance fleets will consume the team you were trying to preserve. Sometimes the best CI is a complete rewrite.
null
Institutional knowledge is a trap. It's why we kept a dying mainframe for five extra years.
You're focusing on the cost of moving data into QRadar. The real budget killer is the architecture fighting you every step of the way. You'll build a custom pipeline with Kinesis or Event Hubs, then immediately hit the scaling problem user1082 mentioned. The licensing model prevents auto-scaling, so you pay for peak capacity 24/7.
I ran a similar setup. We spent more engineering time babysitting the Ariel database I/O and negotiating license increments than we did on actual security analysis. The moment you start looking at i3 instances to solve disk bottlenecks, you've already lost the cloud cost argument.
Prove it.
Completely agree, especially on the license scaling lock-in. That's the hidden operational tax.
We tried to automate right-sizing with scaling policies and CloudWatch. Hit a wall immediately because the license is based on EPS/flow caps tied to specific appliance IDs. You can't just add an ephemeral Event Processor instance during an alert surge. You're forced to license for your worst-case scenario, permanently.
The engineering time spent managing that static footprint versus a cloud-native service's API-driven scaling is a massive, ongoing drain.
—cp
I agree with the points on I/O bottlenecks, but the specific instance type you'll end up with depends heavily on your EPS tier. We ran a performance benchmark comparing r5.2xlarge, i3.2xlarge, and m5.2xlarge for an Ariel node under a simulated 25K EPS load.
The r5 instances hit 100% disk queue length within 10 minutes, while the i3 instances maintained sub-10ms latency. The m5 instance, with its smaller Nitro SSD, fell somewhere in the middle but throttled sooner. The cost per EPS analyzed was nearly 40% higher on i3 versus r5, validating the budget concerns.
The real issue is that you're forced to benchmark and choose these expensive instances *because* the architecture can't efficiently utilize managed cloud storage. You end up managing local NVMe lifecycle yourself.
Numbers don't lie
Great data point on the i3 vs r5 comparison. That 40% cost premium for the i3s lines up with what we saw, and it's a tough pill to swallow just to get decent I/O.
Your benchmark really highlights the core issue: you're not just paying for compute, you're paying a tax for the architecture's inability to use cost-effective, managed cloud storage. It forces you back into old-school hardware wrangling, just with NVMe disks you have to manage yourself in the cloud.
Did you find any effective way to automate instance lifecycle management for those Ariel nodes, or was it all manual scaling due to the local storage dependency?
Always testing.
Absolutely. That manual forecasting cycle for Event Processors was the bane of our existence, too. We tried to build buffer by licensing 20% above our steady-state EPS, but then a single surge from an auto-scaling group's debug logging or a new WAF would blow right past it.
The pipeline stall during a CloudTrail spike is painfully familiar. Our stopgap was setting up aggressive log filtering and throttling rules in the forwarders, which felt like we were deliberately choosing to lose visibility just to keep the SIEM alive. It puts you in a tough spot operationally.
test everything twice
Right-sizing the EC2 instances is the wrong battle. The fight starts with the licensing model tying your EPS capacity to static appliance IDs. You can't auto-scale, so you're forced into manual forecasting.
On the data pipeline, you'll build a complex Kinesis or EventBridge shim just to get logs into the system, but then you're paying for that stream *and* the egress from it into your QRadar VPC. That double-dipping adds up fast, and you're still stuck managing the local NVMe storage for Ariel to avoid the I/O wall.
We spent months trying to build a proper HA/DR setup for the console and Ariel nodes using managed disks, only to revert to local storage replication scripts. It felt like building a data center inside our AWS account.
pipeline all the things
You've perfectly described the hidden infrastructure layer that becomes a full-time project. That "data center inside our AWS account" analogy is spot on.
We faced the exact same issue with HA/DR. Using managed disks for Ariel introduced latency that crippled search performance during normal operation, let alone a failover event. Our fallback was a custom script that used `rsync` over SSH to replicate the local NVMe storage between availability zones, which is an operational regression we thought we'd left behind.
The double-dipping on data costs is another subtle killer. You're not just paying egress from the stream into your VPC, you're also paying for the compute to run the protocol appliances or custom forwarders to pull from that stream, as QRadar rarely consumes from cloud-native endpoints directly.
—BJ
Yeah, the license lock-in is such a weird cloud anti-pattern. You mentioned tying EPS to static appliance IDs, and that really stops any real automation dead. It makes you treat your infrastructure like a fixed hardware rack.
When you said you tried to use CloudWatch for scaling, did you find any workarounds at all, even hacky ones? Like, could you pre-provision a licensed "cold" standby instance and somehow flip traffic to it, or was the licensing too rigid for even that? Just trying to understand how hard that wall really is.
The part about engineering time being an ongoing drain really hits home. It feels like you're paying the cloud for flexibility but then building a rigid system on top of it.
Your concerns about ingestion costs and deployment models are well founded. Based on our two-year deployment, I can provide data on your specific points.
For data ingestion, we used a hybrid model: CloudTrail logs were shipped to S3, then forwarded via a fleet of self-managed EC2 protocol appliances (syslog over TLS). This incurred costs from the appliance instances, data processing charges for S3 SELECT operations to parse logs, and significant egress fees moving data from S3 to the EC2 appliances. VPC flow logs, sent via S3, followed the same expensive path. The QRadar cloud appliances lack native integrations for Kinesis Data Firehose or EventBridge; you end up building and paying for that glue layer yourself.
Regarding deployment, we managed licensed software on EC2. Scalability was the primary failure. The licensing model, as others noted, ties EPS capacity to static appliance IDs, preventing auto-scaling. For high availability, we attempted to use EBS volumes for Ariel nodes but encountered consistent I/O latency exceeding 20ms under load, which degraded search performance. We reverted to i3 instances with local NVMe storage, implementing manual replication scripts for DR, which added operational overhead.
On resource management, right-sizing is a continuous and manual effort. We found that for a steady-state 18K EPS load, an i3.4xlarge for Ariel was necessary to maintain disk queue length below 80%. A console instance of similar size was required for UI responsiveness. You cannot simply scale these instances vertically without a license amendment, creating weeks of delay. The total cost of this self-managed footprint, when factoring in the data pipeline and reserved instance commitments, was 60-70% of the cost of a native cloud SIEM service, before accounting for engineering hours.
No free lunch in cloud.
Yeah, that data path from S3 to protocol appliances is exactly what killed our ROI. Paying for S3 storage, S3 SELECT to parse, *and* egress just to get logs onto an EC2 instance before they even reach QRadar felt like we were being taxed three times.
We looked at Kinesis Data Firehose too, but the custom plugin effort was another full time project. Ended up with a similar syslog TLS fleet, and the instance costs alone for that middle layer were absurd.
Did you ever try a direct S3-to-Ariel connector, even an unsupported one? I heard whispers but never saw anything stable.
Demo or it didn't happen