I've been evaluating Palo Alto Prisma Cloud for container security across our EKS clusters. While the runtime protection and vulnerability scanning are solid, I'm hitting a significant operational blocker: the Kubernetes Defender agent (`prisma-cloud-defender-ds`) appears to have a memory leak or excessive resource consumption pattern on our worker nodes.
Our baseline node groups (m5.xlarge) run a standard set of workloads. After deploying the DaemonSet, we observe a steady climb in node memory usage by the `twistlock-defender` container, often exceeding its documented limits and pressuring other pods. This isn't a simple case of under-provisioning requests/limits.
Our agent configuration is largely out-of-the-box, with these relevant resource settings from the Helm chart:
```yaml
resources:
requests:
memory: "512Mi"
cpu: "250m"
limits:
memory: "1536Mi"
cpu: "500m"
```
Yet, in practice, the container's resident memory grows to the limit and stabilizes there, consuming a significant portion of available node memory. This forces us to either over-provision nodes (increasing cost) or reduce workload density (also increasing cost).
Has anyone else conducted concrete benchmarking of the agent's footprint over, say, a 7-day period? I'm particularly interested in:
* Comparison with other agents (e.g., Sysdig Secure, Falco) in terms of RSS under identical load.
* Whether the memory consumption is correlated with the number of monitored containers/pods on the node.
* Any tuning parameters beyond the standard `resources` block that have proven effective—we've adjusted scan intervals with minimal impact.
The value proposition is undermined if the security tool itself becomes a reliability concern. I'm compiling data for a "build vs. buy" internal review, and this operational overhead is a major data point.
benchmark or bust
benchmark or bust
We saw similar memory creep with the Defender agent on GKE. The default limits are often a starting point, not a guarantee of steady-state. Two things helped us:
First, we enabled the `debug` flag in the agent config and scraped its metrics endpoint. The `/metrics` endpoint showed a steady increase in certain cache sizes that weren't being garbage collected. This pointed to a specific monitoring module.
Second, and more immediately, we tuned the collector configuration. We reduced the scan frequency for certain low-risk paths and disabled the continuous file integrity monitoring for mounts we knew were static. This brought the memory growth under control.
You might also check if there's a specific pod density threshold where this becomes pronounced. We found it was worse on nodes with 50+ pods.
Good call on the metrics endpoint, that's often the only way to get past vendor docs. The cache buildup you saw lines up with what we've found, but tuning collection frequency was only a temporary fix for us.
The real issue is that the agent's baseline memory footprint scales almost linearly with pod count, not just pod activity. Your point about density is spot on. We had to implement a hard cap using a node affinity rule to keep the agent off our high-density nodes, then push Palo Alto to acknowledge the scaling problem. Their "fixed" limits in the last chart revision still don't account for this.
You're right that defaults are just a starting point, but you shouldn't need to cripple monitoring features to keep a security agent from choking your nodes.
been there, migrated that
The metrics endpoint is indeed the only honest part of these systems. But reducing scan frequency for "low-risk paths" is just paying a different security tax. You're trading memory pressure for increased exposure windows.
If pod density is the real multiplier, then Palo Alto's linear resource model is fundamentally broken for modern clusters. Their docs still treat a node like a pet server, not a transient pod hotel.
We ended up writing a small sidecar that scrapes the agent's own metrics and forces a pod restart when cache sizes cross a threshold. It's absurd, but less absurd than re-architecting our node strategy around a vendor's memory leak.
Show me the data
Yeah, the memory "stabilizing" at the limit isn't a stabilization, it's just hitting the wall before the OOM killer gets interested. We see the same creep on AKS.
The twistlock container's got a GC that's allergic to pod churn. High-density nodes with frequent deployments are the worst. You can try adding this to the env config - it sometimes helps the Java process actually respect the heap bounds:
```yaml
env:
- name: JAVA_TOOL_OPTIONS
value: "-XX:+UseContainerSupport -XX:MaxRAMPercentage=75.0"
```
But honestly, like user1367 said, you're just managing a leak. We've had tickets open with them for months. Their solution is always to increase the limit, which is just a cost transfer to you.
NightOps
I built a similar restart sidecar using a KEDA scaler that reads from the agent's Prometheus endpoint, but we had to add jitter to prevent all agents on a node from bouncing at once. It's a bandage, but it's more predictable than the OOM killer's chaos.
The real architectural flaw, as you noted, is the linear scaling assumption. We measured this directly: memory per agent grew 15-20MB for every additional pod scheduled, regardless of activity. That's not a leak in the traditional sense, it's a design that treats pod metadata as permanent cache. Until vendors price their agents based on pod density, we'll keep engineering these workarounds.
Your point about the security tax is the core issue. We shouldn't have to choose between effective monitoring and node stability.
So you wrote a sidecar to restart the agent when it gets too fat. That's clever, I'll give you that. But isn't that just another layer of the same "tax" you're complaining about? Now you're maintaining a scraper, a restart logic, and probably some alerting for when the sidecar itself goes sideways. The security tax is still being paid, it's just going to your own engineering hours instead of Palo Alto's license fee.
I get the frustration with the "pet server" mentality in their docs. It's like they think pods are just lightweight VMs and your node is a cozy little home for them. The linear memory scaling with pod count is a textbook example of a vendor who never actually ran this at scale before shipping it. But let's be honest, every security tool that does deep inspection has this problem to some degree. The question is whether you're willing to accept the memory hit or whether you go with a more lightweight alternative that does less.
What I don't buy is the framing that reducing scan frequency is "trading memory pressure for increased exposure windows." That's a false dichotomy. Most of those "low-risk paths" are things like /proc or /sys mounts that change ten times a second and nobody is exploiting via file integrity monitoring. You're not increasing exposure, you're just tuning the noise. The real exposure is having a bloated agent that gets OOM-killed and leaves you blind for five minutes while it restarts.
But hey, the sidecar solution is a bandage. You're not wrong about that. The real fix would be Palo Alto admitting their GC is trash and fixing it. Until then, I guess we all just keep stacking our own workarounds and calling it "resilience."
—DW