Skip to content
Notifications
Clear all

Just built an automated response to quarantine pods flagged by Sysdig.

2 Posts
2 Users
0 Reactions
30 Views
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
Topic starter   [#10867]

Sysdig's runtime alerts are great, but manual pod quarantine is a bottleneck. Built an automated responder using their webhooks and a small Kubernetes operator.

Logic is straightforward:
1. Sysdig webhook fires on a `Runtime Container Drift` or `Unexpected Process` alert with high severity.
2. Simple service receives the webhook, validates the payload.
3. Operator patches the offending pod with a `quarantine` label.
4. NetworkPolicy (already deployed) denies all ingress/egress traffic to any pod with the `quarantine` label.

Key part of the operator's reconciliation logic:

```go
func (r *PodQuarantineReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
pod := &corev1.Pod{}
err := r.Get(ctx, req.NamespacedName, pod)
if err != nil {
return ctrl.Result{}, client.IgnoreNotFound(err)
}

if pod.Labels["quarantine"] == "true" {
// Check for existing restrictive network policy
if !hasQuarantineNetPol(pod) {
// Apply deny-all NetworkPolicy to pod's namespace targeting label
err = r.applyNetworkPolicy(ctx, pod)
if err != nil {
return ctrl.Result{}, err
}
}
}
return ctrl.Result{}, nil
}
```

Benefits:
* Stops lateral movement instantly.
* Pod stays up for forensics.
* Logs all actions for audit.

Pitfalls:
* Needs tight control on what triggers the webhook to avoid false positives.
* Operator requires RBAC to patch pods and create NetworkPolicies.
* Sysdig alert payload must be sanitized to prevent K8s object injection.

Works as a stopgap before a full SOAR integration. Considering open-sourcing the core.

-dk


Trust but verify, then don't trust.


   
Quote
(@harryj)
Reputable Member
Joined: 3 months ago
Posts: 381
 

Nice approach! The webhook-to-label flow is solid. Have you considered adding a short delay or confirmation step? We got burned once by a Sysdig alert that fired on a transient deployment artifact.

Could be worth logging the pod's owner (Deployment/StatefulSet) in your service. That way you can notify the owning team automatically, not just isolate and forget.


Automate the boring stuff.


   
ReplyQuote