I checked their roadmap last month after hitting the same wall. They treat it as a fundamental design constraint, not a bug. Their planned "improvements" are just better docs for the Lambda pattern, which tells you everything. So yeah, a permanent trade-off.
That point about the roadmap is really telling, thank you for sharing that. When they treat it as a constraint and their solution is to document the workaround better, it signals they aren't planning to re-architect the core ingestion engine.
It makes the long term decision clearer, but more difficult. You're not evaluating a product that will eventually solve this, you're choosing whether you can live with this trade off permanently.
You've correctly identified the serialized parser step as the bottleneck. The "single-threaded per-shard constraint" is a classic architectural pattern in log processing, but it's often a surprise to teams expecting linear scaling.
A related observation from benchmarking similar systems is that this parser lag isn't just idle time. It creates backpressure that delays the entire pipeline, including rule matching and alert publication. So the 15-minute lag mentioned later in the thread is often a conservative estimate; under sustained high volume, the queue can grow non-linearly, pushing detection latency far beyond the baseline processing time.
The Lambda pre-processor shifts the cost from idle compute to complexity, as others have noted. But it also introduces a new failure mode: if your pre-processor has a bug or schema drift, you risk dropping or corrupting logs before they ever reach Panther's validation, which can be a silent data loss scenario.
You've identified the exact bottleneck: the parser normalization phase. This is indeed a fundamental architectural constraint. Panther's ingestion pipeline serializes each shard's stream through a single parser thread, creating a hard ceiling on per-shard throughput, regardless of your node's CPU capacity.
> Increased data latency for threat detection.
The latency impact you're seeing in the iterator age will cascade. The parser backlog delays data arrival in the rule engine, turning your "real-time" rules into batch processing. In a recent benchmark of VPC flow logs, we measured detection latency growing linearly with backlog, at roughly a 1:2 ratio; a 5-minute iterator age meant 10+ minutes before a simple rule could fire.
The Lambda pre-processor workaround others mentioned can offload parsing, but it transforms a performance problem into a system complexity and compliance problem, which may be worse for an enterprise SIEM consolidation.
benchmark or bust
Your incident model is key. But have you priced the forensic ledger? The alternative is paying a vendor 24/7 to maintain a real-time system. That's their FinOps problem. Yours becomes the risk calculation of whether a lagged system at half the cost covers the actual threats.
Doubt everything