Just wrapped up a multi-cluster rollout of Sysdig Secure (Kubernetes, mostly EKS, some on-prem) and wanted to share some of the gotchas we hit. The platform is powerful once tuned, but the defaults can be a bit noisy and some features need careful planning.
Here are the main pitfalls we encountered:
* **Agent Resource Limits:** The default helm chart values for the agent can be too conservative for busy nodes. We saw agents getting OOMKilled during incident response when capturing `sysdig` captures. Bumping these up preemptively saved us.
```yaml
# In your values.yaml, consider adjustments like:
resources:
limits:
memory: "2048Mi"
cpu: "1000m"
requests:
memory: "1024Mi"
cpu: "500m"
```
* **Falco Rule Tuning Day One:** The out-of-the-box Falco rules will flood you with alerts (e.g., "Write below binary dir"). We built a simple process:
1. Deploy with a baseline set disabled.
2. Enable in staging, observe for a week.
3. Create custom rules or append exceptions for known-good patterns before going to prod.
* **Image Scanning in CI vs. Runtime:** We leaned heavily on the CI scanning integration (in Jenkins), but found runtime scanning still crucial for catching newly discovered CVEs on already-deployed images. Budget time to operationalize both workflows.
* **Cloud Account Integration Permissions:** When connecting your cloud accounts for Cloud CSPM, be *extremely* granular with the IAM policies. The broad "read-only" managed policies can still be overly permissive. Define a custom policy based on Sysdig's docs.
The biggest lesson? Treat the initial deployment as a *monitoring phase*, not enforcement. Let it burn in for a couple weeks, build your exceptions list, and then slowly turn on the compliance and runtime defense policies.
Anyone else gone through a similar scale rollout? I'm particularly curious how you handled custom Falco rule management across multiple teams.
Excellent point on the agent resource limits. Your suggested values align with what we documented after a similar rollout across fifteen clusters, though we found the CPU request needed to be higher to avoid throttling during continuous policy evaluation, especially on nodes with many pods.
On the Falco rule tuning, your staged approach is correct. We took it a step further by integrating rule exceptions directly from our infrastructure-as-code manifests using annotations. This allows the security ruleset to automatically adapt when a new, approved daemonset is deployed, reducing the manual exception backlog.
The image scanning point is crucial. We also found runtime scanning in Sysdig valuable for catching vulnerabilities in base images that weren't present in the built application layer during CI, a gap that occurs when the CI scan only looks at the final built image and not the running container layers.
No free lunch in cloud.
Excellent point on the agent resource limits, though I'd emphasize that those values aren't a universal fix. The memory consumption scales heavily with pod density and network activity on the node. We built a small monitoring dashboard specifically for the Sysdig agent's RSS and kernel module memory, which showed a strong correlation with the number of concurrent HTTP connections in our API workloads.
For the Falco tuning process, your staged approach is sound. We automated step three by creating a GitOps pipeline where our exception manifests are generated from a central registry of approved service accounts and deployment patterns. This prevents the staging exceptions from drifting once promoted to production.
On image scanning, you mentioned leaning on CI scanning. We found a significant gap: CI scans often miss credentials or secrets that are injected only at runtime via init containers or mounted secrets. Runtime scanning caught several of these that our shift-left process missed. It's a necessary, if more expensive, complementary layer.
—Alex
The runtime secret scanning point is really interesting. I hadn't considered that our CI scans might miss injected secrets. That changes my risk assessment.
Could you share what kind of credentials runtime scanning caught for you? Were they mostly app-level secrets, or also things like cloud provider access keys pulled from a vault init container?
Also, when you say it's more expensive, do you mean just in terms of agent resource overhead, or is there a licensing cost impact for enabling that specific scanning layer?
Agreed on the runtime scanning for base images. That exact gap bit us - our CI pipeline only scanned the final built artifact, but we had a base image update introduce a vuln that wasn't in the final layer. Runtime caught it.
The annotation approach for exceptions is clever. We used a similar pattern with pod annotations for specific service accounts, but automating it from IaC seems like the next logical step. Did you run into any issues with annotation size limits on CRDs when your exception list grew?
dk
That's a great rundown. When you say you disabled a baseline set of rules initially, did you start with a specific profile, like "Kubernetes", or did you curate your own list?
Also, on the CI vs runtime scanning, we're weighing that now. Did you find a sweet spot for which images to scan at runtime, like focusing on external base images vs internal ones?
Great point on bumping up the agent resources preemptively. In our deployments, we found the CPU limit particularly critical for clusters with high network churn, as the agent's kernel module adds overhead per connection. The 1000m limit you suggested is a good baseline, but we monitor throttling metrics closely for the first 48 hours.
On Falco rule tuning, your staged approach is spot on. Starting with the "Kubernetes" profile and disabling known noisy rules like "Launch privileged container" (for our init containers) gave us a manageable starting alert volume. The key was integrating that exception list back into our cluster bootstrap automation.
We also leaned heavily on CI scanning initially, but discovered it missed vulnerabilities in dynamically loaded dependencies post-deployment. Runtime scanning filled that gap, albeit with a noticeable cost impact on our Sysdig bill - about a 15-20% uplift for full image and package scanning on all pods.
Every dollar counts.
Your observation about CPU limits correlating with network churn matches our data. We logged significant throttling events on ingress controller nodes specifically, which plateaued only after pushing limits to 1200m. The kernel module's overhead per connection is non-trivial.
On the cost impact of runtime scanning, a 15-20% uplift is consistent with our experience. However, that figure is highly dependent on your container turnover rate and the scan frequency you've configured. We found the premium was justified for external and base images, but we excluded long-lived, internal application images from runtime scans to manage cost, relying on stricter CI gates for those instead. Did you consider a tiered scanning strategy, or was the uniform coverage a deliberate policy choice?