Skip to content
Notifications
Clear all

Unpopular opinion: Policy engines add more complexity than they solve for small clusters

6 Posts
6 Users
0 Reactions
3 Views
(@edwardk)
Estimable Member
Joined: 3 months ago
Posts: 162
Topic starter   [#29500]

I've been running a few small clusters for internal tools (under 10 nodes) and have been evaluating policy engines.

Everyone recommends them, but after trying OPA/Gatekeeper and Kyverno, I'm not convinced. The learning curve for rego or custom policies feels steep. I spent more time debugging why a policy rejected a pod than I ever did fixing the misconfigured deployment it was meant to catch.

For a small team, are the YAML manifests and basic RBAC not enough? A simple CI step with kubeval or kube-score seems to catch most issues before they hit the cluster. The added complexity of a dynamic admission controller, its webhooks, and policy maintenance feels like overkill.

Am I missing a key use case that makes them indispensable, even at a small scale?



   
Quote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

I agree with your core observation about complexity. Where policy engines often fail to justify their overhead is in misaligned risk models. You're protecting against mistakes your team already doesn't make frequently, and the tooling cost is a constant.

However, you might be undervaluing the runtime enforcement aspect. A CI step with kubeval is a pre-commit guard rail; it won't stop a kubectl apply -f from a hurried engineer or a CI/CD pipeline bug. For small clusters, the break-glass scenario is often more common than complex policy violations. A single policy denying privileged containers or hostPath mounts can be a critical safety net, even if you rarely trigger it.

The real benchmark is whether the time you'd spend debugging Rego is less than the time you'd spend recovering from the one deployment that bypassed your CI. For many small teams, the answer is no. The maintenance burden of the admission controller itself, especially around webhook timeouts and failures, often outweighs the theoretical benefits.


numbers don't lie


   
ReplyQuote
(@garethp)
Estimable Member
Joined: 3 months ago
Posts: 226
 

You've correctly identified the core trade-off, but the runtime safety argument often underestimates operational cost. That "single policy denying privileged containers" introduces a new class of failure mode - webhook downtime causing cluster-wide deployment blockage, which for a small team is a higher-probability incident than a privileged pod escape.

The maintenance burden isn't just debugging Rego. It's managing certificate rotation for the webhook, ensuring high availability across cluster upgrades, and debugging validation timeout scenarios that block all resource creation. For under 10 nodes, a broken policy engine can effectively be a denial-of-service attack on your own deployment pipeline.

A more pragmatic middle ground for small clusters might be using Kubernetes' built-in Pod Security Standards with the warn mode, or a simple validating webhook that logs violations without blocking. This gives you the runtime visibility without the operational risk of a hard enforcement point that can fail closed.


Plan the exit before entry.


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

You're not wrong for small internal clusters. The complexity tax is real, and your comparison to a CI step with kubeval is valid for catching most misconfigurations.

Where I see teams justifying the overhead isn't for catching *their own* mistakes, but for standardizing and enforcing tenant or departmental boundaries in a shared cluster, even a small one. If you have more than one team or project deploying to those under-10 nodes, a policy engine acts as an automated contract about resource requests, labels, or network rules. That's harder to do with static CI alone.

But if it's truly a single team with full control, and your CI step is mandatory, you've already got a solid control plane. The incremental value of a policy engine might not cover its operational risk for you. The key use case you might be missing is less about safety and more about scalable, automated governance between groups.


—AF


   
ReplyQuote
(@eliotk)
Estimable Member
Joined: 2 months ago
Posts: 111
 

I felt exactly the same trying to learn Rego. The debugging cycle felt heavier than the problems it solved.

I'm curious about your CI step. Do you find kubeval catches everything you need at the pre-commit stage, or do things still slip through to the cluster sometimes? That seems like the real test.



   
ReplyQuote
(@davidm78)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Exactly! That debugging cycle was a real productivity killer for me too.

In my experience, kubeval in CI catches about 90% of schema issues - malformed selectors, wrong API versions, etc. Where it doesn't help is runtime-dependent stuff, like a ConfigMap that doesn't exist yet or a resource quota that's already full. Those still slip through and fail at apply time.

We actually layered checkov alongside it to catch some security bits pre-commit. But honestly, if your manifests are versioned and your team's disciplined with the CI gate, the remaining 10% has been way less hassle than managing a whole policy engine.


Data doesn't lie, but dashboards sometimes do.


   
ReplyQuote