Skip to content
Notifications
Clear all

What actually works for runtime security in a 5-eng team?

3 Posts
3 Users
0 Reactions
25 Views
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
Topic starter   [#16801]

Hey folks! Our small team is trying to get serious about runtime security for our k8s workloads. We're all-in on GitOps with Argo CD, and our CI is GitHub Actions.

We need something that actually works without becoming a full-time job. I'm curious what others are doing at this scale. Specifically:
* Does it integrate with our pull request workflow? Can we fail a PR if a new image has critical CVEs?
* Can we define policies as code (maybe in the repo next to our k8s manifests)?
* How noisy is it? We can't have 100 alerts a day.

I've been looking at Aqua's inline scanner in CI. Something like this in a GitHub Actions step seems promising:

```yaml
- name: Scan image
uses: aquasecurity/trivy-action@master
with:
image-ref: 'myapp:${{ github.sha }}'
format: 'sarif'
output: 'trivy-results.sarif'
```

But what about *runtime*? Do we need a daemonset, or can we start simpler? What's been your experience? 😅

> git commit -m 'done'


git push and pray


   
Quote
(@chrism)
Reputable Member
Joined: 3 months ago
Posts: 326
 

Yeah, that's exactly how we started - Trivy in GitHub Actions for image scanning at PR time. It's great for catching CVEs before they hit the cluster.

For runtime, we skipped the daemonset at first and went with Trivy's Kubernetes operator. It scans workloads *after* they're deployed by Argo CD, which felt simpler than managing another daemonset. It can post findings back to Slack and, crucially, you can set policies right in the repo. We have a ConfigMap with policies that blocks deployments with critical CVEs, and Argo just won't sync them. The noise was high until we tuned the policies to ignore low/unfixed stuff.

Honestly, that combo - scanning in CI for the gate, and the operator for runtime drift - has covered maybe 90% of what we need without a huge overhead. You could even start with just the CI scans and add runtime later.


K8s enthusiast


   
ReplyQuote
(@cloud_ops_amy_2)
Reputable Member
Joined: 7 months ago
Posts: 274
 

That's a solid approach. We did the same thing but found the operator's runtime scans could lag a bit, especially with ephemeral dev namespaces. We ended up adding a scheduled scan job in CI that runs weekly against our production image tags and posts a cost-of-fixing report.

One caveat on the policy ConfigMap: if you manage it via Argo too, watch out for circular sync issues. We keep ours in a separate Argo app that syncs first.


terraform and chill


   
ReplyQuote