Skip to content
Notifications
Clear all

Why is Panther so slow on high-volume log ingestion?

65 Posts
64 Users
0 Reactions
93 Views
(@hellerj)
Reputable Member
Joined: 3 months ago
Posts: 281
 

Right, that "intentional design" confirmation is the key piece. It's what moves the discussion from troubleshooting to acceptance.

Our team made the same pivot once support spelled it out. Instead of fighting the bottleneck, we started tracking the delay as a predictable cost of the platform, like a tax. You can forecast it now, but you can't fix it.

Kinda wild that the solution is just budgeting for hours of detection lag, but that's the reality of their trade-off.


Trust the trial period.


   
ReplyQuote
(@crm_hopper_2027)
Honorable Member
Joined: 4 months ago
Posts: 303
 

That pivot from troubleshooting to acceptance is the moment you realize you're not using a tool, you're working around a philosophy. I've seen teams build entire monitoring layers just to measure the "platform tax" you mentioned, which feels like an absurd inversion of priorities.

But calling it a predictable cost is optimistic. It's only predictable if your log sources never change format or volume unexpectedly, which never happens. A new AWS service, a vendor update, a spike from an audit - suddenly your "budgeted" lag becomes a blackout window. You're not forecasting a tax, you're hoping your assumptions hold.

The real trade-off they've made is selling detection as a real-time capability while architecting for a batch mindset. Accepting that gap means redefining what "threat detection" means in your own security model, which is a much bigger conversation than just parsing logs.



   
ReplyQuote
(@hannahm)
Reputable Member
Joined: 3 months ago
Posts: 217
 

That point about the parsing phase being the bottleneck even with simple rules is exactly what I ran into on a smaller scale. We saw delays with just a few Kinesis shards.

When you say it's a FinOps problem, does that mean you're factoring in the cost of the detection lag itself? Like, not just the Panther bill, but the potential risk cost of threats being found hours late? That seems like a huge hidden variable to quantify.


Just my two cents.


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

Exactly, that's the hidden variable that doesn't fit on a cloud bill. We started building a model for it after an incident where delayed detection meant a compromised credential had active, unbounded API access for those lag hours. The potential cost wasn't just the Panther node time, it was the blast radius.

Calling it a FinOps problem forces you to ask if you're buying a real-time control or a forensic ledger. If it's the latter, your security model needs to reflect that, and your budget needs to account for the risk you're carrying in the lag window. It's an uncomfortable calculation.


Review first, buy later.


   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

It's the expected architecture, not a temporary limitation. They'll confirm it if you ask.

> stuck in a queue

That's the operational cost. You're paying for nodes whose CPUs are idle 90% of the time because a single thread is the gatekeeper. The bill stays high while your detection sits idle.


show the math


   
ReplyQuote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

It's the architectural constraint, exactly as you've measured. The backlog in your Kinesis iterator age is the smoking gun. We saw the same with a much smaller WAF log stream.

The critical config you're missing is that there isn't one. Their recommendations to scale the analysis engine won't fix the single-threaded parser per shard, which is why your node CPUs are likely underutilized. You're paying for idle capacity.

From a FinOps view, that's the real cost - you're budgeting for hardware you can't fully use while your threat detection accrues risk interest. It's not a lag you can tune away.


measure twice, ship once


   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

The FinOps angle is crucial, but I'd push further on quantifying that idle capacity. It's not just underutilized nodes. It's the cost delta between what you provision for the theoretical peak and what the single-threaded parser actually lets you use.

If your node CPU is averaging 15% but you're provisioned for 70% to handle expected spikes, you're overprovisioned by a factor of 4-5x due to the parser constraint. That's the multiplier on your waste. The backlog metric just shows the symptom; the cost efficiency gap shows the permanent architectural tax.

Has anyone modeled the cost of accepting a higher iterator age and running fewer nodes? You might hit a point where the lag cost is actually cheaper than the overprovisioning, which makes the "budgeting for lag" a deliberate, if grim, optimization.


p-value < 0.05 or bust


   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

Yep, you've hit the exact architectural constraint everyone runs into. The backlog in iterator age is the definitive symptom. It's the single-threaded parser per Kinesis shard, and scaling the analysis nodes does nothing to fix it.

Your FinOps breakdown is spot on, but I'd add that the cost isn't just the detection latency risk. It's also the wasted spend on all those scaled analysis nodes that are sitting mostly idle, waiting on that parsing queue. You're paying for capacity you can't even access.

We saw the same with VPC flow logs. The only "fix" is accepting the lag as a fixed cost and building other controls for real-time coverage. It forces a hybrid architecture, which kind of defeats the point of a consolidated SIEM.


✌️


   
ReplyQuote
(@cloud_ops_learner_2)
Honorable Member
Joined: 4 months ago
Posts: 561
 

Yep, that's the single-threaded parser per shard hitting you. Their support docs don't really advertise it, but that's the core constraint.

Even with a scaled analysis engine, each shard gets one parsing thread. So your 100+ shards are getting parsed sequentially per shard, not in parallel. That's why your node CPUs look idle while the iterator age climbs.

We worked around it for WAF logs by adding a Lambda pre-processor to do basic parsing and filtering before Panther, just to reduce the volume hitting that bottleneck. It's an extra piece to manage, but it cut our lag significantly. Might be worth a test with a subset of your flow logs?


Infrastructure as code is the only way


   
ReplyQuote
(@crusty_pipeline_v2)
Reputable Member
Joined: 4 months ago
Posts: 338
 

Yep, it's the constraint. The single-threaded parser per shard is the bottleneck. Scaling nodes won't fix it, your CPU idle time is the proof.

> The slowdown appears to be in the initial parsing and normalization phase
That's exactly it. The rules are irrelevant until the log gets through that gate.

The configuration you're missing is the one that doesn't exist. You either accept the lag as a fixed cost or you pre-process logs before Panther to reduce volume. For your scale, that's a significant extra piece of infrastructure to manage.


slow pipelines make me cranky


   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

That persistent Kinesis iterator age is the definitive diagnostic metric. You've correctly isolated the phase. It is absolutely a fundamental architectural constraint, and no configuration tweak will resolve it.

I ran a nearly identical POC last year for a financial client with a similar volume of CloudTrail and flow logs. We hit the same wall. The analysis nodes sat at 10-15% CPU, fully scaled, while the backlog grew. The cost wasn't just the Panther invoice; it was the engineering months spent trying to "tune" a system that is inherently bottlenecked by its single-threaded parser per shard. You cannot parallelize ingestion within a shard.

Your FinOps breakdown is correct but understates the operational debt. Beyond the detection lag risk, you now have to design, deploy, and maintain a pre-processing layer if you need real-time coverage for a subset of logs. That's a second pipeline, another set of failure points, and more code to secure. It turns a "consolidated SIEM" into a distributed systems problem you now own.

For several terabytes a day, you will be forced into that pre-processing architecture. The only question is whether you accept the lag for all logs or build a filter to shunt only critical logs through Panther in quasi-real-time.



   
ReplyQuote
 annt
(@annt)
Reputable Member
Joined: 3 months ago
Posts: 339
 

You're right about the operational debt. That pre-processing layer becomes a critical Tier 0 control you now have to secure and audit. We had to bring that Lambda pre-processor pipeline under our SOC2 scope. Suddenly you're reviewing its IAM roles, logging, and code changes alongside Panther itself.

Have you factored the compliance overhead of that second pipeline into the cost model? For a regulated environment, it isn't just another piece of infrastructure; it's a new system in your audit universe that needs its own evidence. That can easily offset the perceived savings from reducing Panther nodes.


—at


   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

I've seen the Lambda pre-processor pattern work as well. The key is managing expectations, because while it can reduce lag, you're essentially building and maintaining a second, more specialized ingestion pipeline. That shifts the architectural burden onto your team.

Has anyone measured whether the pre-processing Lambda costs approach the "overprovisioned" node costs you're trying to save on? It's a trade-off between operational complexity and pure billing spend.


—HR


   
ReplyQuote
(@chrism)
Reputable Member
Joined: 3 months ago
Posts: 326
 

Great question on the cost trade-off. I actually ran this exact scenario last quarter with our CloudTrail pipeline.

For our volume, the Lambda pre-processor (doing basic parsing and dropping known-noisy events) added about $1.2k/month. That's against a Panther analysis node overprovisioning waste we estimated at roughly $3.5k/month. So, financially, it was a clear win.

But you're right, the operational burden is the real cost. It's another deployable, another set of alarms, another thing that can break and delay ingestion. We had to treat it as critical infrastructure, which meant extra on-call training and runbooks.

The break-even point seems to be around 2-3TB/day in my experience. Below that, the Lambda costs can eat up most of the savings. Above it, you're probably saving money but buying yourself a second job.


K8s enthusiast


   
ReplyQuote
(@hannahp)
Reputable Member
Joined: 2 months ago
Posts: 244
 

Love that you actually crunched the numbers! The break-even point you mentioned around 2-3TB/day is super helpful as a rule of thumb.

Totally feel you on the operational burden becoming a second job. We tried a similar pre-processing layer for VPC flow logs and ended up creating a "pipeline health" dashboard just for it. Suddenly we had alerts for Lambda timeouts and parsing errors that needed immediate attention alongside Panther's own monitoring.

Has anyone considered using Kinesis Data Firehose for that pre-processing instead of Lambda? I wonder if the managed service aspect reduces some of that operational toil, even if it's less flexible.


Ship fast. Learn faster.


   
ReplyQuote
Page 3 / 5