We're a 50-engineer product team on AWS. We need a SIEM for security compliance (SOC 2, etc.) and to monitor our own AWS accounts, containers, and apps. We've tried the big names (Splunk, Datadog) and the bill is a joke for our data volume.
Primary requirements:
* Ingest ~100 GB/day of mixed logs (CloudTrail, VPC Flow, app logs).
* Must handle multi-account AWS structure.
* Team needs to build detections without a dedicated security engineer.
Panther is on the shortlist because it's AWS-native and promises lower costs. Need real numbers and operational overhead.
Specific questions for those running Panther at scale:
1. **Actual Cost:** What's your monthly AWS bill breakdown for Panther? Don't talk list prices. I want S3, Dynamo, Lambda, and OpenSearch costs.
2. **Management Burden:** How many FTE hours per week to keep it running? Our team can't babysit it.
3. **Detection Scaling:** Does the Python-based rule engine work for 50+ active contributors, or does it become a governance nightmare?
If you switched from another SIEM, what was the real cost delta? Show me the math.
cost per transaction is the only metric
Right there with you on the sticker shock from the big guys. Our Panther setup for about 80GB/day runs around $2.8k/month on AWS, mostly OpenSearch. S3 is cheap storage, Lambda is pennies. The big variable is your OpenSearch instance sizing for that 100GB/day retention.
Management is maybe 2-3 hours a week once it's stable. The initial multi-account setup is the heavy lift. After that, it's mostly updating detections.
The Python rules are a double-edged sword. For 50 engineers, you need a solid PR review process in a single repo from day one, or it gets messy fast. It's powerful but requires discipline.
We cut our Datadog security bill by about 65%. The math was painfully clear.
Always optimizing.
Glad the math worked for you, but I'm suspicious of that "2-3 hours a week once stable" line. That only holds if your team's log sources never change and your detections are static. Add a new service, tweak a parsing rule, and suddenly you're debugging why half your alerts went silent.
And that 65% savings against Datadog... were you comparing apples to apples? Panther's cheaper because you're the one building and maintaining the logic. Add the eng time for those Python PR reviews, and the gap narrows fast.
Trust but verify.
That 2-3 hour maintenance estimate is only valid for a truly static environment, which almost never exists. The operational cost isn't just rule updates; it's the regression risk when you modify the underlying data pipeline or a log format changes.
You'll need to budget for a weekly sync, maybe 30 minutes, to review Panther's own health metrics (failed ingestions, rule errors) and at least a few hours quarterly for version upgrades. The multi-account structure adds fragility; a single misconfigured S3 bucket policy in a spoke account can silently stop a log source.
On the cost delta, we saved about 60% vs. Splunk, but the real math should be: (Previous Vendor Cost) - (Panther AWS Cost + (Estimated Eng Hours * Your Fully Loaded Rate)). For a 50-engineer team, even a 5% time allocation for governance and troubleshooting eats heavily into the raw infrastructure savings. The value is in control, not just cheaper storage.
Your real cost question is the right one, but you're still thinking about it wrong.
> Don't talk list prices. I want S3, Dynamo, Lambda, and OpenSearch costs.
That's the easy part. For 100GB/day, your AWS bill will likely be $3-4k, mostly OpenSearch. The brutal math is the denominator. You said your team needs to build detections without a dedicated security engineer. That's your biggest cost.
If 50 engineers can write Python detections, you need:
* A strict, enforced schema for all logs (who's building that?)
* A CI/CD pipeline for rule testing and deployment (that's a project)
* A librarian to manage the rule repo and prevent breaking changes
Other replies say 2-3 hours a week. That's fantasy for a dynamic team. It's more like one engineer-week to get it stable, then 5-10 hours a week of collective time reviewing PRs, debugging silent alerts, and fixing ingestion breaks. At a fully loaded rate, that's another $2-4k/month easily.
The cost delta looks good until you realize you've just hired a part-time, junior security engineer spread across 50 people. It works if you treat it like any other critical platform service, with ownership and SLAs. If you treat it as a side task, it will fail silently and you'll miss your SOC 2 controls.
You've nailed the real cost that everyone else is tap-dancing around: the denominator in that equation is always a fantasy. "Estimated Eng Hours * Your Fully Loaded Rate" is pure fiction until you've lived through a breaking log format change that your detection library depends on.
The silent log source failure you mentioned is the perfect example. That's not a 30-minute weekly sync to fix, that's a post-incident retrospective where you discover your compliance coverage has been a lie for three weeks. The control is valuable, but it's bought with constant, paranoid vigilance, not a fixed hourly budget.
And let's be honest, for a 50-engineer team without a dedicated security person, who exactly owns that vigilance? You'll either create a thankless security-champion rotation that becomes a career dead end, or it'll fall to the platform team as yet another "undifferentiated heavy lifting." The savings are real, but they're a transfer of risk, not an elimination of it.
Trust but verify.
Thanks for laying out the numbers you're after. We're looking at Panther too, and our biggest hang-up is the rule governance question.
> Does the Python-based rule engine work for 50+ active contributors, or does it become a governance nightmare?
This is my fear exactly. Everyone talks about the AWS bill, but what's the process cost? If 50 engineers can commit Python, you need a rock-solid schema for logs and a mandatory CI gate that actually blocks bad rules. Who enforces that on a team with no dedicated security person?
Did anyone build a lightweight review process that actually worked without becoming a full-time job?
Still learning
> The cost delta looks good until you realize you've just hired a part-time, junior security engineer spread across 50 people.
This is the most useful way I've seen this framed, thanks. You're saying the operational cost isn't just hours, it's the organizational burden of coordinating 50 people.
But is there a middle ground? Could you start by locking down rule writing to a small group, maybe 2-3 engineers, and let the other 47 just submit detection ideas as tickets? That seems like it could keep the governance nightmare manageable while still using the team for threat modeling.
Containers are magic, but I want to know how the magic works.
That middle ground you propose is precisely the model that fails in practice. You're describing a two-tier system where a small group of gatekeepers manages the logic, and the broader team submits tickets. The governance overhead doesn't vanish, it simply shifts.
The "detection ideas as tickets" become a backlog requiring triage, prioritization, and translation from a vague request into a precise, testable Python rule. Those 2-3 engineers now need deep context on every service's log schema and threat model. They become bottlenecks, and the ticket queue grows stale, defeating the purpose of leveraging 50 engineers' domain knowledge.
A more functional approach is to invert the model: enforce a rigid, versioned schema contract for all log sources at the point of emission. If the schema is immutable and validated on ingest, then any engineer can write a rule against a known, stable interface. The governance moves upstream to the schema definition, not the rule writing. This is harder to establish initially but scales. Without that enforced contract, your small core team will drown in parsing exceptions.
You're absolutely right about shifting the governance burden upstream to schema definition. It's the only scalable model. However, that initial contract is a massive political and technical lift for a team without a dedicated security function.
I've seen teams try to enforce that immutable schema with a centralized "log library" published to internal package registries. It creates a hard, versioned API boundary. The break happens when a service team needs to add a new field for debugging and can't wait for a library release cycle, so they just tack it onto the JSON. Suddenly your "immutable" contract is fiction.
The real failure mode isn't the core team drowning in parsing exceptions, it's them being circumvented entirely.
—Alex
You've put a fine point on the real cost. I'd add that your formula is missing a term for unplanned work.
> (Previous Vendor Cost) - (Panther AWS Cost + (Estimated Eng Hours * Your Fully Loaded Rate))
The "Estimated Eng Hours" is a best-case scenario. The real cost includes incident response hours when a detection you didn't know was broken fails to fire. That's not a predictable weekly maintenance task; it's a random, high-context interrupt that burns senior engineer time.
Your point about silent log source failures is exactly this. The cost isn't just fixing the S3 policy. It's the investigation into what you missed during the outage window. That's where the governance time isn't just an allocation; it's a risk buffer you're forced to carry.
Data is the only truth.
Yeah, that two-tier model feels right in theory, and I've seen teams try it. The snag is that the translation layer from "detection idea" to working Python rule requires so much specific context about log schemas and threat models that the gatekeepers become a major bottleneck.
Those tickets don't just get written, they need clarifying questions, back-and-forth, and testing against real data. So the 2-3 engineers you've appointed become full-time security rule developers, not just reviewers. You're right back to a dedicated resource, just under a different title.
Maybe a better middle ground is to let everyone write *tests* for rules, but the core logic is still owned by a small team. That way you're crowdsourcing validation of the logic against their own services without the free-for-all on production code.
ship it
Yes, the schema contract is the critical pivot point. The problem is that you've just moved the bottleneck upstream to the architecture review process for new log sources. Who approves that versioned schema? Without a dedicated owner, you're relying on voluntary compliance during design phases, which is the first thing that gets dropped under deadline pressure.
I've seen teams enforce this by mandating that all log producers ship with a JSON Schema or Protobuf descriptor to a registry, and the SIEM ingestion pipeline rejects anything not matching a registered version. This works, but it's a significant CI/CD and culture change. It's less about the technical implementation and more about getting 50 engineers to treat their audit logs as a serious, versioned API.
Logs don't lie.
You've perfectly described the exact failure path I've witnessed twice now. The centralized library works until the pressure to deliver a feature hits. The debugging field is the classic Trojan horse, a seemingly harmless addition that bypasses the entire governance model.
What made it stick on the third attempt was pairing the schema registry with a mandatory, automated enrichment stage in the pipeline itself. Any log event missing a registered schema version gets routed not to the SIEM, but to a low-latency stream the developer can see in their own staging environment. The pain of their detection not working becomes immediate and personal, not a ticket for the central team. It turns a governance problem into a self-service debugging issue.
The key was making the circumvention more painful than compliance, but without adding to the central team's burden. They don't own the broken stream, the service team does.
This is a fascinating solution, and I agree that shifting the cost of non-compliance directly onto the service team is the only sustainable model. The crucial detail is that "low-latency stream the developer can see." It fails if that feedback loop takes minutes or hours.
You need the validation and routing to happen at the edge, ideally in the same function or sidecar emitting the log. If a developer pushes a schema-breaking change, they should see the detection stream go empty in their next automated integration test, not hours later when they check a dashboard. This makes the failure a build-time concern, not an operational one.
My caveat is the overhead of maintaining that staging or test stream infrastructure. It requires a parallel, high-fidelity data pipeline that mirrors production routing but with different sinks, which itself becomes a distributed systems problem to keep reliable and performant.
brianh