Skip to content
Notifications
Clear all

Has anyone benchmarked Panther's query performance on 10TB+ of data?

5 Posts
5 Users
0 Reactions
24 Views
(@cost_analyst_liam)
Honorable Member
Joined: 6 months ago
Posts: 515
Topic starter   [#19551]

Having recently undertaken a comparative analysis of SIEM platforms for a client with a substantial on-premises data lake, I found a distinct lack of granular, cost-attributed performance data for Panther at petabyte scale. While the marketing materials discuss "sub-second" queries and "high-performance" data lakes, these terms are operationally meaningless without a defined architectural context and associated expense profile.

My primary concerns, which I hope other members with large-scale deployments can address, revolve around the intersection of performance and the underlying cloud vendor billing mechanics. Specifically:

* **Query Engine Resource Allocation:** Panther's default use of Presto/Trino is clear, but the cost driver is the size and auto-scaling configuration of the virtual warehouse. For ad-hoc searches over 10TB of log data:
* What is the typical node instance type (e.g., AWS m5.8xlarge, Azure D32_v5) and cluster size required for a sub-30-second response on a complex JOIN across 30 days?
* How aggressive is the auto-scaling down, and what is the observed "cool-down" period where costs are incurred for idle resources? This directly impacts the unit economics per query.

* **Data Layout & Hidden Storage Operations:** Performance is dictated by data organization. The documentation mentions automatic partitioning.
* Is this partitioning applied to the underlying S3/GCS objects, or is it a metastore-level abstraction? Mismatches here lead to expensive full scans.
* Are there background optimization jobs (e.g., compaction, re-partitioning) that incur recurring compute and API request costs (like S3 LIST, GET)? These often appear as line items separate from the core platform cost.

* **Network Transfer & Egress Fees:** A 10TB dataset implies significant data movement. When the query engine reads from object storage and returns results to the UI:
* What is the typical data processed/returned ratio for a standard detection search? Transferring 1TB of scanned data to return 1KB of alerts is a severe cost inefficiency.
* In a multi-cloud or multi-region setup, are there configurations that inadvertently cause cross-region data transfers, invoking the punitive egress fees of cloud providers?

I am seeking empirical observations, not theoretical architecture. If you have run Panther at this scale, please share the cloud provider, the annual/log data volume, the typical monthly cost breakdown for compute (warehouse), storage (including API calls), and data transfer. Details on any required tuning to avoid performance cliffs that correspond with billing spikes would be particularly valuable.

-- Liam


Always check the data transfer costs.


   
Quote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

You're asking the right questions. Those "sub-second" claims always assume you've thrown a lot of money at the cluster size. I've seen the default configs on mid-sized deployments, and they're often undersized for the JOINs you're talking about. You'll be waiting minutes, not seconds, unless you manually crank it up.

The real kicker is the auto-scaling "cool-down." In practice, it's never as aggressive as the sales deck implies. You're paying for those big nodes to sit there for way longer than you'd like, because spinning them up and down adds latency they don't want to risk. So your unit cost per query gets bloated with idle time. It's the oldest cloud trick in the book.

Anyone giving you performance numbers without the associated hourly cloud bill is just telling you a fairy tale.


Your stack is too complicated.


   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

You've nailed the core issue: the marketing terms are useless without the cost context. Your point about the node instance type is critical - everyone asks about size, but the instance family (memory optimized vs compute optimized) and network bandwidth can be a bigger determinant for those large JOINs than just core count.

I can't give you exact numbers for a 10TB JOIN, but I've seen similar workloads choke on the default 'general purpose' instances because of spill to disk. The config that finally worked was on AWS with r-type nodes (r6g.8xlarge) to keep everything in memory, and even then you're looking at a 20+ node cluster for that data volume. Sub-30 seconds is possible, but you're paying for that memory premium constantly because of the cool-down lag.

Have you looked into forcing a specific warehouse configuration via their APIs? It's a pain, but you might claw back some control over the scaling policy that way.


pipeline all the things


   
ReplyQuote
(@claraj)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Exactly. The moment you're talking about r-type nodes, you're paying the AWS tax twice. First for the overprovisioned memory, then for Panther's markup on top.

Their API for forcing a warehouse config is a red herring. It might let you specify a node size, but the auto-scaling cooldown is still baked into their control plane. You'll just be locking in a more expensive idle state.

So you get sub-30 seconds for a 10TB JOIN, but only while your $30k/month cluster is already warmed up and idle. That's not a benchmark, it's a confession.


Prove it


   
ReplyQuote
(@ellaj8)
Reputable Member
Joined: 3 months ago
Posts: 295
 

You've pinpointed the exact problem. Those performance figures are always divorced from the billing reality.

I can't give you node counts for a 10TB JOIN, because it's the wrong metric. The real question is your concurrency. If you're the only one running a massive query, you can force a huge, memory-optimized cluster and get your sub-30 seconds. The cost per query will be absurd, but you'll get it. The moment you have two analysts running concurrent heavy searches, everything changes. Now you're into queueing, spill to disk, and minutes of wait time.

The auto-scaling cooldown is the silent killer. They talk about scaling to zero, but in practice the control plane nodes have a minimum runtime measured in hours, not minutes. You're paying for the 'scaffolding' even when the workers are gone. So your unit economics are anchored to that idle baseline.

A better benchmark would be: "What's the hourly cost to sustain a 95th percentile query latency of 45 seconds with three concurrent users on 10TB of data?" You won't get that from Panther. You'll get it from your cloud bill after a month of painful tuning.


Trust but verify – and audit


   
ReplyQuote