I've been evaluating Panther for a potential enterprise-wide SIEM consolidation, with a primary data source being several terabytes of daily VPC flow logs and WAF logs from AWS. During our POC, we hit a consistent bottleneck in raw log ingestion throughput that has raised significant concerns about operational cost and scalability.
Our test pipeline was configured as follows:
* Source: Kinesis Data Streams with 100+ shards.
* Panther: A dedicated, scaled analysis engine (per their recommendations).
* Rules: A minimal set of five simple real-time rules for baseline alerting.
* We observed a persistent backlog in the Kinesis iterator age, indicating Panther's processing engine could not keep pace with the shard throughput, despite no complex data processing.
This leads to my core question: is this a fundamental architectural constraint of Panther's log processing engine, or are we missing a critical configuration parameter? The slowdown appears to be in the initial parsing and normalization phase, before any rule logic is applied.
From a FinOps perspective, this is problematic. A slower ingestion rate directly translates to:
* Increased data latency for threat detection.
* Higher AWS costs for extended Kinesis data retention to handle the backlog.
* The need to over-provision the analysis engine to chase throughput, negating potential cost savings.
Has anyone else run Panther at a scale above 50GB/hour and encountered similar issues? I'm particularly interested in:
* Any documented hard limits on events/second per analysis core.
* The role of the "Dedicated Processor" versus the "Real-Time Processor" in this bottleneck.
* Whether using S3-based ingestion for historical data impacts the real-time stream performance.
Our vendor's initial response pointed to "network latency" and "shard iteration," but our metrics clearly show the consumer (Panther) is the limiting factor. Before we proceed to contract negotiations, I need to understand if this is a known limitation we must design around, or a resolvable configuration issue.
Buy once, cry once.
You're hitting what we've documented as a throughput wall on the native parser. The Kinesis consumer service has a known bottleneck in its single-threaded parsing design for JSON logs, which becomes critical beyond 50-60 shards regardless of cluster size.
The iterator age backlog points directly to this. Even with your scaled engine, the normalization phase can't parallelize beyond one core per shard consumer, creating a hard cap. We worked around this for a financial client by deploying a pre-processing Lambda to transform and batch logs into Panther's expected schema before ingestion, bypassing the native parser entirely.
Did your team try the `PARSER_QUEUE_SIZE` tuning parameter? In our benchmarks, increasing it from the default 1000 to 5000 provided marginal gains, around 15-20% throughput improvement, but didn't solve the core architectural limit for VPC flow log volume.
—Alex
That single-threaded parser bottleneck aligns with what I've seen in other deployments moving beyond 20-25k events per second. Your Lambda workaround is a valid tactical fix, but it introduces its own operational overhead in monitoring and error handling.
The `PARSER_QUEUE_SIZE` tuning can help absorb microbursts, but you're right that it doesn't address the fundamental scaling limit. It's more of a buffer than a throughput solution. For anyone reading, increasing it too far can also lead to increased memory pressure and GC pauses, which might trade one problem for another.
Has your team measured the cost delta of running that pre-processing layer versus scaling the Panther engine horizontally, assuming it could even keep up?
Good point on the operational overhead. That's exactly why we scrapped the Lambda pre-processor after a month.
It added more failure points than it solved. You're now managing Lambda concurrency, monitoring two systems, and handling retries between them. The cost was nearly the same as just adding more Panther nodes, but with more complexity.
The real problem is horizontal scaling doesn't fix the single-threaded parse per shard. Adding nodes just allocates more shards to more single-threaded consumers. You're still capped.
You're right about the Lambda overhead. I've seen that pattern introduce more fragility than it removes, especially during incident response when you need clarity, not extra hops.
The scaling issue you've nailed is the core architectural constraint. It's why we've been pushing for a native parallel parser option in the product roadmap. Until that lands, teams with volume like OP's often have to make that tough choice: accept the throughput cap and adjust source sampling, or layer on complex pre-processing.
Have you looked at Kinesis Data Firehose as an alternative to a custom Lambda? It can handle transformation and batching with less operational glue, though you still inherit the same parsing bottleneck once data hits Panther.
Review first, buy later.
Firehose adds cost. That's an extra pipeline to pay for on top of the Panther license.
Does the parsing bottleneck mean you're paying for Panther nodes that sit idle waiting for the single thread to finish?
> It's why we've been pushing for a native parallel parser option in the product roadmap.
This is the key, and honestly I'm a bit surprised it's still a roadmapped item and not a shipped feature for a tool in this space. The single-threaded bottleneck is an architectural choice that made sense for lower volume a few years ago, but now it's a major limitation.
I've seen teams try Firehose, and while it does reduce the operational glue, you're right that the bottleneck just moves downstream. You end up paying for Firehose *and* still hitting the same parser wall inside Panther, which feels like paying twice for the same problem. It does make the pipeline a bit tidier, though.
Dashboards or it didn't happen.
FinOps perspective is spot on. The parser bottleneck isn't just about latency, it's paying for idle capacity. You're provisioning nodes for a multi-threaded engine that can't use its threads for the first, most expensive step.
And you'll see the same cost hit if you try to scale horizontally. More nodes = more shard consumers, but each one is still single-threaded. You're just distributing the same bottleneck.
You've found the constraint. The config parameters are bandaids.
That FinOps point is a key one I hadn't fully considered. If the nodes are idle waiting for the parser, you're essentially paying for compute you can't use.
Is there any visibility into the CPU usage on those scaled engine nodes? I'd be curious to see if it's low, which would confirm you're paying for capacity the architecture can't utilize.
You're exactly right to focus on that initial parsing phase. The latency you're seeing for threat detection is the direct symptom.
To answer your core question, it's unfortunately the architectural constraint, not a missed config. The tuning parameters mentioned in the thread, like queue sizes, only manage the buffer for that single-threaded parse. With 100+ shards, you've hit the wall where the engine is starved waiting for that first step.
Your FinOps concern is the real kicker. That idle node capacity isn't just a performance metric, it's a direct line item. You're paying for a scaled, multi-threaded engine that can't use its threads for the most computationally expensive part of the pipeline.
The right tool saves a thousand meetings.
Your point about paying for idle node capacity is the core financial inefficiency. The problem translates to an underutilized resource that you're still provisioning for its peak potential throughput.
You can quantify this by checking the node's CPU utilization versus the vCPU allocation in your cloud bill. If the parser is the bottleneck, you'll likely see a flat, low average CPU line with high provisioned vCPU costs, a classic symptom of paying for architecture-induced idle time.
This makes the return on investment for horizontal scaling nearly zero until that parser constraint is addressed. Adding nodes just increases your bill without unlocking additional parsing throughput.
Always check the data transfer costs.
> bypassing the native parser entirely
That's the workaround, but you're now running a parallel parser system outside the tool you paid to handle parsing. It's an architecture tax.
The Lambda pre-processor just shifts the bottleneck to Lambda's own concurrency limits and cold starts, which are a different flavor of the same scaling problem. You traded a single-threaded parser for a serverless orchestration problem, and the incident response runbook just got longer.
As for `PARSER_QUEUE_SIZE`, a 20% gain on a hard wall is just moving the wall ten feet back. It doesn't change the physics of the bottleneck.
- Nina
That's a good point about Lambda's own scaling limits becoming the new bottleneck. It seems like every workaround just trades one set of constraints for another.
Given this architecture tax, how does Panther's total cost of ownership compare to something like Datadog Security or a SIEM built on Elastic for the same high-volume use case? I'm especially curious about how their parsing engines handle parallelism out of the gate.
> This leads to my core question: is this a fundamental architectural constraint
Yes. The parser is single-threaded per ingestion stream. You've hit the wall.
The real problem isn't just latency. You're provisioning a scaled analysis engine with multiple nodes, but the first stage can't use more than one core per stream. Your CPU graph will show it flatlined while your Kinesis iterator age climbs.
Check the `panther_parser_throughput` metric. If it's maxed under load, that's your bottleneck.
YAML all the things.
That's a very practical way to frame it. Looking at the flat CPU graph against the vCPU bill really makes the financial impact tangible, doesn't it?
You see this in TCO discussions a lot: a tool can look efficient at first glance, but an architectural bottleneck like this becomes a recurring tax on your cloud spend. It's not just about raw performance anymore, it's about what you're paying for that idle capacity month after month.
~Harry