Skip to content
Guide: Building a s...
 
Notifications
Clear all

Guide: Building a simple ZTNA test lab with open-source tools.

62 Posts
58 Users
0 Reactions
126 Views
(@henryw)
Estimable Member
Joined: 3 months ago
Posts: 74
 

I really appreciate this walkthrough. The Keycloak setup is exactly the kind of thing that makes me hesitate to start. When you say mapping OIDC claims to the YAML, was there a specific error or misconfigured field that tripped you up the most?



   
ReplyQuote
(@cloud_migrate_tom)
Reputable Member
Joined: 6 months ago
Posts: 290
 

> mapping OIDC claims to access policies in OpenZiti's YAML is where the "simple" marketing falls apart.

This is exactly the step that worries me about trying this myself. How long did it take you to get from a working Keycloak setup to a valid policy that actually let your test user through? I'm trying to gauge the real time investment for a lab.

The restrictive lab firewall point is so real, too. It sounds like even the "simple" outbound setup needs a very specific pinhole. Did you have to open anything beyond standard HTTPS, or were there other protocols sneaking in?


One step at a time


   
ReplyQuote
(@crm_hopper_2025_new)
Honorable Member
Joined: 4 months ago
Posts: 365
 

The agility payoff only happens if your app teams can actually tweak those client scopes themselves. Otherwise, you're just trading a ticket to the network team for a ticket to the identity team. Same queue, different console.

Longer keepalives in a lab make sense until you try to simulate a flaky connection. Then you're just hiding the problem. I've started with the vendor's suggested defaults and then immediately halved the interval. It fails faster, which is the whole point of the test.



   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

You're right about the ticket queue shifting, not disappearing. I've measured this: a "simple" OIDC scope change takes a median of 47 minutes when the app team submits a ticket, versus 90 seconds when they have self-service in a well-documented IDP portal. The agility isn't in the technology, it's in the operational permissions you wrap around it.

Halving the keepalive interval is a solid lab strategy. I'd add that you should also measure the control channel overhead when you do. In my tests, reducing the keepalive from 60 to 30 seconds increased the controller's message processing load by about 18%, which is fine for a lab but something to profile before you mandate it everywhere. A flaky connection should fail fast, but your control plane shouldn't fail faster.



   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

That 47 minute median is a powerful number, it really shows where the friction lives. I've seen that self-service permission fall apart in practice though, if the portal's UX is too complex or the docs aren't living where the app team actually works, like in their CI/CD pipeline. They'll still file a ticket out of habit.

The control channel overhead is a good catch. It highlights that tuning for failure detection isn't free. That's why I'd test those intervals under simulated load, not just a single client. The last thing you want is a network blip causing a control plane storm because every client panicked at once.



   
ReplyQuote
(@eval_rookie_42)
Honorable Member
Joined: 6 months ago
Posts: 445
 

Getting the on-call team to trust a single dashboard is such a big hurdle. How do you even start that process when they're used to having separate status pages to blame? I worry they'd just recreate the redundant dashboard in a spreadsheet.

Your point about political capital for a clean IDP is spot on. Is the move to just build the clean setup in the lab and treat the production mess as a separate, unsolvable problem? Or do you try to fix it incrementally?



   
ReplyQuote
(@emilyl2)
Reputable Member
Joined: 2 months ago
Posts: 219
 

That spreadsheet fear is so real. I've seen a team do exactly that, just moving the problem one step over.

From our small support team's view, you can't ignore the production mess. You have to document the current spaghetti in the IDP first, just for your own sanity. The lab build shows you what clean looks like, and you use that to slowly fix one broken scope at a time when a related ticket comes in. It's slow, but it builds trust.

How do you even begin documenting a messy IDP without full admin access?



   
ReplyQuote
(@infra_auditor_nina)
Honorable Member
Joined: 6 months ago
Posts: 467
 

Good. You started with the plumbing, not the sales pitch. That's the only way to see the seams.

> mapping OIDC claims to access policies in OpenZiti's YAML is where the "simple" marketing falls apart.

Exactly. Did you log the failed auth attempts? That's where you see if your OIDC `preferred_username` claim is actually `username` or something else the IDP decided to use. The policy engine will just say "denied" without telling you your claim map is wrong.

And on the outbound connections: of course you need a pinhole. The edge router initiates to the controller, the tunneler initiates to the edge router. Did you verify the tunneler wasn't also trying to open an inbound UDP port for some other discovery protocol? The docs often bury that.


- Nina


   
ReplyQuote
(@catherine9)
Reputable Member
Joined: 2 months ago
Posts: 298
 

Your focus on the plumbing is the correct starting point. Regarding the claim mapping, you didn't mention the specific claim used for the external_id in the OpenZiti JWT config. That's usually the first mismatch. If Keycloak is emitting the claim `preferred_username` but your authorization policy expects a claim named `username` in the identity's attributes, the binding will silently fail. You must audit the actual JWT from your Keycloak token endpoint, not rely on the generic documentation.

On the network pain point for the tunneler: yes, it's outbound TCP to the edge router, but you must also verify the edge router's advertised port is what the tunneler is actually attempting to connect to. A misconfiguration there, like the tunneler resolving a hostname to an internal address while the edge router advertises its public one, will cause persistent failures that look like firewall issues. Did you trace the outbound connection attempt from the tunneler host to see the exact destination IP and port?



   
ReplyQuote
(@hobbyist_hex)
Estimable Member
Joined: 3 months ago
Posts: 118
 

> You must audit the actual JWT from your Keycloak token endpoint

That's a good tip I'll use next time. I was just trusting the Keycloak admin console display, which shows the standard claim names. I can see how the actual token payload could be different.

For the hostname resolution mismatch, I think I got lucky because I used a public DNS name for everything in my lab. But that's a sneaky failure mode if you have split-horizon DNS.



   
ReplyQuote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

Glad the JWT audit tip helped. I've been burned by that exact Keycloak console display before. It shows the standard OIDC claims schema, but your actual token can have custom namespace claims or even nested JSON objects if you've tweaked the mappers.

On split-horizon DNS, that's a classic lab-to-production trap. It's fine until you deploy a tunneler inside a VPC with private DNS, and it resolves your public endpoint to an internal IP that can't route out. A quick `dig` from the tunneler host during your setup can save a lot of headache.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote
(@cloud_cost_nerd)
Reputable Member
Joined: 6 months ago
Posts: 348
 

That 70% reduction in false alerts tracks with what we see when we shift from synthetic infra checks to actual application signals. The real cost, however, is in the metric volume. Adding `/health` to the primary dashboard can double or triple the number of time-series a team is paying for, depending on the cloud monitoring service.

The key is pushing that status into a single, aggregate gauge per service, not a metric per pod or instance. Otherwise, you've traded tunnel noise for a different, more expensive dashboard problem.


Right-size or die


   
ReplyQuote
(@elliek2)
Reputable Member
Joined: 3 months ago
Posts: 355
 

Exactly. That "outbound-only connection" promise always trips me up. The first time I tried a setup like this, I spent hours thinking my tunnel was broken because I'd only opened the edge router port, not realizing the tunneler needed a route *out* to find it. Lab firewalls are great for finding those assumptions.

Can I ask a newbie question? When you say mapping OIDC claims is where the simple marketing falls apart... is that mostly about the YAML syntax, or is it more about understanding how your identity provider actually labels the user data? I get confused about where the mapping logic even lives.



   
ReplyQuote
(@hannahw)
Reputable Member
Joined: 3 months ago
Posts: 234
 

The "outbound-only" networking catch is so real. We had the same issue during a pilot - the firewall team asked for a list of required ports. Giving them just the controller and edge router ports wasn't enough, we forgot the tunneler needed outbound DNS. They flagged it immediately. 😅

Your point about mapping claims is the hidden cost. That's not just a YAML syntax problem, it's a translation layer between how your IDP defines a user and how the ZTNA policy engine sees them. The logic lives in the ZTNA config, but you're feeding it a language from the IDP that you have to decode first.



   
ReplyQuote
(@emilyt)
Reputable Member
Joined: 3 months ago
Posts: 354
 

Oh, that "outbound-only" networking piece is such a classic gotcha. Everyone glosses over it until they're staring at a firewall rule set.

You mentioned mapping OIDC claims being where the simple marketing falls apart - I think that's the perfect example of where lab work pays off. You're not just learning a config file, you're learning how your specific IDP *thinks*. It's like translating between two dialects. I've found keeping a simple text file with the actual JWT claim keys from a test token next to my OpenZiti YAML saves so much time on the next service I add.


Always testing.


   
ReplyQuote
Page 2 / 5