Hey folks, been running OPA/Gatekeeper for cluster policy enforcement for about two years. Made the jump to Kyverno on our new EKS clusters last quarter. Wanted to share the "why" and, honestly, some of the friction we hitβbecause no migration is ever perfectly smooth 😅
For us, the main drivers were:
* **Native K8s style:** Writing policies as Kubernetes resources just clicked better with the team vs. Rego. It lowered the barrier for our platform engineers who aren't full-time policy language devs.
* **Mutation & Generation:** This was the killer feature. Being able to mutate non-compliant resources (e.g., adding standard labels) or generate new ones (like NetworkPolicies) automatically saved so many dev tickets.
* **Context-aware rules:** Validating across resources felt more straightforward. For example, ensuring an Ingress has a corresponding Service is simpler to express.
But it wasn't all sunshine. The pain points we felt:
* **Policy Exceptions:** We missed OPA's `dry-run` feature for testing. Kyverno's `PolicyException` CRD is powerful, but managing a separate resource for one-off exemptions feels heavier.
* **Scale & Performance:** At very high resource churn (like in our CI namespaces), we saw a noticeable increase in admission latency compared to our tuned Gatekeeper setup. We're still tweaking the Kyverno config.
* **Debugging:** When a complex policy fails, the error messages are sometimes less informative than we'd like. Figuring out *why* a validate rule denied something can take a bit of digging in the logs.
Overall, the trade-off was worth it for our use caseβthe team adoption and the mutation features outweighed the cons. For anyone else considering it, I'd say if your team is deep into Rego and needs extreme performance, maybe stick with OPA. But if you want something that feels more "Kubernetes-native" and can leverage mutations, Kyverno is a fantastic choice.
Curious if others have hit similar issues, or found clever workarounds for the exception management?
✌️
βοΈ