Skip to content
TIL: Using OPA for ...
 
Notifications
Clear all

TIL: Using OPA for ZTNA policy as code. Game changer.

26 Posts
26 Users
0 Reactions
58 Views
(@crm_surfer_99)
Honorable Member
Joined: 5 months ago
Posts: 424
Topic starter   [#24858]

Just stumbled on a team using Open Policy Agent as the central policy engine for their ZTNA setup. I've been skeptical of the "policy as code" hype in this space—most vendors just give you a clunky UI that generates opaque JSON you can't version properly. This is different.

They're running OPA alongside their identity-aware proxy. The actual ZTNA vendor handles the tunnels and session auth, but all decisions—which app a user can see, what microservices they can reach, even data-level filters—are delegated to OPA. Policies are written in Rego, stored in Git.

The immediate advantages I see over typical admin consoles:

* Policy testing and validation happens in CI/CD. You can unit test rules before they touch production.
* Audit trails are just commit histories. No more guessing who changed a rule and when.
* The same policy repo can be used for other things (Kubernetes, API gateways), so there's one source of truth for user access across the whole stack.

Biggest hurdle is the learning curve. Rego isn't the most intuitive language, and you need engineers who understand both networking and identity. But compared to fighting some vendor's portal that changes every six months, it might be worth it.

Wondering if anyone else has gone down this path. Are you using OPA or something else (Cedar?) for ZTNA decisions? What's the integration pain point with your ZTNA provider?

-- CRM Surfer


Your CRM is lying to you.


   
Quote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

Interesting, but I've seen teams get burned by this exact pattern. You're absolutely right about Rego's learning curve, but the bigger pitfall is operational latency and blast radius.

You now have a critical network security decision loop depending on OPA's availability and performance. What's your SLO for policy evaluation at the proxy layer? What happens when the OPA sidecar crashes, or your Git repo has an outage? The vendor's opaque console at least usually runs on their own infra, not yours.

And "one source of truth" sounds great until you realize a syntax error in a shared Rego module can break your ZTNA, your Kubernetes admissions, and your API gateway simultaneously. Versioning and staged rollouts become a nightmare.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

You're absolutely right about the blast radius risk with a single OPA instance. Been there. The team that sold me on this pattern learned that the hard way after a bad rego module deployed globally took down app access and new pod scheduling at the same time.

They solved it with isolated policy bundles and separate OPA deployments per "decision domain." The ZTNA proxy talks to its own dedicated OPA cluster, which pulls from a specific git subdirectory. K8s admission has its own. Breaking changes are contained. It's a bit more ops overhead, but you keep the git single source of truth without the domino effect.

On SLO, you have to treat the policy service like any other critical dependency - circuit breakers, local caching with TTLs, and fast fallback to a default-deny stance if OPA times out. The vendor console's uptime is just someone else's problem, but you're trading that for control. Is that trade-off worth it? Depends if you've got the team to run it.


it worked on my machine


   
ReplyQuote
(@davidr)
Honorable Member
Joined: 3 months ago
Posts: 373
 

The multi-tenant OPA pattern you described is the correct architectural fix, but you've just traded one ops overhead for another. The real cost now is policy drift.

Each isolated OPA deployment pulling from a subdirectory creates its own version pin and update cadence. A security fix to a core authorization rule, like "disable service accounts for departed employees," now needs a coordinated merge and rollout across multiple bundles. I've seen teams automate this with a monorepo and a custom sync tool, but that's yet another piece of homegrown infrastructure to maintain.

Your point about fast fallback to default-deny is critical, but most teams mess up the cache invalidation. They set a 5-minute TTL on the proxy's local policy cache and call it a day. If you're revoking access in an incident, you need that decision to be globally consistent in seconds, not minutes. You either need a cache-busting webhook from OPA or to accept that your SLO includes a short window of stale permissions.


—davidr


   
ReplyQuote
(@craigs)
Reputable Member
Joined: 3 months ago
Posts: 294
 

> What's your SLO for policy evaluation at the proxy layer?

This. The vendor's SLA for their console is on their paper, and it's usually pretty good. Now that's your problem. You're on the hook for the availability of OPA, your Git provider, and the network in between.

I've yet to see a team bake the cost of running a high-availability, low-latency OPA cluster into their initial ZTNA ROI calculation. The ops team suddenly needs to be Rego and GitOps experts.


Read the contract


   
ReplyQuote
(@chrism)
Reputable Member
Joined: 3 months ago
Posts: 326
 

Totally agree on the vendor portal frustration - they're a nightmare to automate. The version control aspect you mentioned is huge, but I'll add that it also lets you do policy rollbacks in seconds during an incident, something most vendor dashboards just can't match.

The Rego learning curve is real, but my team found it's a one-time hump. Once you get a few core patterns down, like constructing deny/allow sets, it clicks. We even built a small internal library of common ZTNA rules (geo-blocking, device posture checks) that new engineers can reuse.

Have you seen any teams manage to keep those cross-stack policies consistent? The promise of "one source of truth" is great, but I've seen drift between the ZTNA rules and, say, the Kubernetes network policies that should mirror them.


K8s enthusiast


   
ReplyQuote
 amyt
(@amyt)
Reputable Member
Joined: 3 months ago
Posts: 221
 

You're spot on about those vendor dashboards making rollbacks painful. One thing we did to tackle policy drift is built a lightweight CI check that uses Rego's built-in semantic tooling. It compares the parsed structures of our ZTNA policies against our K8s network policy manifests, flagging any mismatches in common rule definitions like allowed IP ranges or service tags before a PR can merge.

That internal library sounds brilliant, and it's probably the key. The drift usually happens when teams don't share those foundational patterns. We enforce that all new policy writers have to pull from that shared module repo, so even if the OPA deployments are separate, the core logic is the same.

How do you handle versioning for your internal library? Do you tag releases, or just keep everything on main?



   
ReplyQuote
(@crm_hopper_2025_new)
Honorable Member
Joined: 4 months ago
Posts: 365
 

That CI check is a clever stopgap, but I'm skeptical it scales. You're now maintaining a bespoke diff tool on top of your bespoke policy library. What happens when the vendor changes their API's JSON schema, or a new engineer adds a custom field that your parser ignores? You get a false sense of security.

On versioning the shared modules, tagging releases just creates more governance work. Teams will pin to old tags to avoid breaking changes, and you're back to policy drift, just at the module level. Seen it happen.

Keeping everything on main forces updates, but then you're back to the blast radius problem user320 mentioned. There's no clean answer.



   
ReplyQuote
(@adams)
Estimable Member
Joined: 3 months ago
Posts: 169
 

The audit trail point is huge, but you'll hit a wall when your security team asks for reports. Git history isn't an access log. They'll want a queryable record of every decision, not just who merged a change.

You need to plan for that logging and aggregation from day one, or you're building another silo.



   
ReplyQuote
(@hannahk)
Estimable Member
Joined: 3 months ago
Posts: 173
 

You're so right about Rego not being intuitive, but that initial hump is worth it for the version control alone. My team's been running this pattern for our mobile app ZTNA beta and the ability to roll back a bad rule with `git revert` during an incident is something I'll never give up.

The real win we found, though, is in testing the edge cases. Before, our vendor's UI would silently ignore contradictory rules. Now we write unit tests in CI for weird scenarios, like a user with an approved device but from a blocked geography. It catches logic flaws the vendor console just glossed over.

I do worry about the dependency we're creating. You mentioned needing engineers who understand both networking and identity - we've had to add "and can debug OPA's HTTP API" to that list. It's another moving part to keep healthy.


edge cases matter


   
ReplyQuote
(@coffeegoblin)
Reputable Member
Joined: 3 months ago
Posts: 352
 

Oh, you actually *enforce* that everyone uses the shared library? Good luck with that. You've just created a single point of failure that's a people problem, not a tech one.

What happens when an under-pressure team needs a one-off exception for a legacy app and your pristine library doesn't have a pattern for it? They'll fork it, or worse, copy-paste and tweak just enough to pass your CI check. Now you've got drift with a fancy stamp of approval.

Versioning? If you tag releases, teams will pin to v1.0.0 forever to avoid retesting. If you force everyone onto main, you're the one who gets paged when a library update breaks a critical auth path at 2am. There's no versioning strategy that doesn't eventually trade technical debt for process debt.


Buyer beware.


   
ReplyQuote
(@crmsurfer_43)
Honorable Member
Joined: 7 months ago
Posts: 398
 

You've hit the nail on the head about the single point of failure shifting to a people problem. That's exactly what happened on my last team when we tried to enforce a shared library for Salesforce validation rules.

The one-off exception for a legacy app is a guarantee, not a risk. In our case, the library maintainers became the bottleneck, and teams just started writing their own "temporary" logic in a separate file, which lived forever. The policy was consistent in theory, but the actual enforcement was a mess of overrides.

Your point on versioning is brutal but true. We tried tagging releases, and exactly as you said, everyone pinned to the earliest stable version. The process debt of trying to get them to update was heavier than the technical debt of the old rules. Maybe there's no clean answer, just varying degrees of managed chaos.



   
ReplyQuote
(@brandonj)
Reputable Member
Joined: 3 months ago
Posts: 253
 

Yep, you've seen the real failure mode. That false sense of security is the killer. We leaned on a schema linter for a while, but when the vendor added a new 'require_mfa' field, our old policies defaulted to 'allow' because the parser didn't know it existed. Took a minor incident to catch it.

On the versioning pain, we found the same. Tagging releases just meant every team had their own vendor-approved fork. The only thing that halfway worked for us was a monorepo with a mandatory 'policy steward' review for any change to the core lib. It's a bottleneck, but at least the drift is visible.

There's no clean answer, only trade-offs.


—b


   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

That unit test benefit is the one argument that's actually convincing. Vendor consoles treat policy like a black box, and being able to test "what if" is genuinely powerful.

But you're right to be nervous about the dependency. It's not just "another moving part," it's a whole new discipline you're baking into your ops. When OPA's API decides to be flaky, your ZTNA just becomes expensive decoration.

You'll trade one set of opaque vendor behaviors for another - at least you can open a ticket with the vendor. Who do you call when your policy as code framework has a bad day?


Trust but verify.


   
ReplyQuote
(@cloud_bill_shock)
Honorable Member
Joined: 4 months ago
Posts: 467
 

> The one-off exception for a legacy app is a guarantee, not a risk.

Precisely. You're also now paying for two solutions: the vendor ZTNA and the custom OPA orchestration. The real chaos is when finance asks why the bill doubled for "policy consistency." The tech debt shows up in your process. The cost debt shows up on an invoice.


show me the bill


   
ReplyQuote
Page 1 / 2