I'm planning a ZTNA pilot for my team and wanted to share my approach to get feedback. We're a mid-size SaaS company with a hybrid cloud footprint (AWS + on-prem legacy apps), and we're looking to move away from our traditional VPN for remote access. The goal is to start small with 10 power users from engineering, security, and IT.
Here's my proposed pilot architecture:
* **Scope:** 5 internal web apps (2 in AWS ECS, 3 on-prem) and SSH access to a subset of development bastion hosts.
* **Identity Foundation:** We already use Okta for SSO, so I plan to leverage that as the primary identity provider. The key will be tying device posture checks (via a lightweight agent) to the user authentication flow.
* **Candidate Solutions:** I'm evaluating a pure cloud service (like ZScaler Private Access or Cloudflare Access) vs. a self-hosted OpenZiti setup. The trade-off is speed vs. control.
* **Success Metrics:**
* User experience score compared to VPN (simple survey)
* Reduction in internal attack surface (measure by closing inbound firewall ports)
* Time to deploy access to a new app for the pilot group
The biggest question I have is about the agent. For this pilot, is it better to go agent-based for the richer context (like verifying disk encryption and OS patches) or agentless for the easiest onboarding? I'm leaning towards a lightweight agent for the power users, as they already have other security tools installed.
Has anyone here run a similar small-scale pilot? I'm particularly interested in how you handled the "split-brain" period where some apps are behind ZTNA and some are still on the legacy network for the same users.
-- Amy
Cloud cost nerd. No, I don't use Reserved Instances.
I'm a security architect at a mid-sized fintech (around 300 employees). We run a hybrid stack similar to yours and migrated from a VPN to ZTNA about 18 months ago; we currently have ZScaler Private Access in production for all corporate users.
Core comparison between a cloud service (ZScaler/Cloudflare) and self-hosted (OpenZiti) for a 10-user pilot:
* **Deployment time and effort**: Cloud services win by a mile. I had our ZScaler pilot (10 users, 3 apps) configured and live in under 2 business days. OpenZiti requires deploying and managing your own controllers, routers, and edge routers; the initial setup alone is a multi-day project. For a pilot, speed is your friend.
* **Initial and ongoing cost structure**: Cloud ZTNA is SaaS-priced. For 10 users, expect a pilot to cost you effectively nothing; list prices for mid-market are typically $7-12/user/month at scale. The hidden cost is the mandatory bundle - vendors often require their full suite. OpenZiti is free software, but your cost is 1-2 FTE weeks to build, maintain, and troubleshoot the overlay network.
* **SSO and posture integration maturity**: With Okta, the cloud services offer pre-built, click-to-configure integrations for both authentication and device posture (using their agents or Jamf/Intune). In our ZScaler pilot, tying Okta groups to app access and requiring a CrowdStrike signal took an afternoon. OpenZiti has the concepts, but you wire most of the logic and policy enforcement yourself - a significant configuration lift.
* **Pilot exit strategy**: This is often overlooked. If the cloud pilot fails, you turn off a config. If the OpenZiti pilot fails, you're decommissioning infrastructure. Conversely, if the cloud pilot succeeds, you're locked into a vendor. With OpenZiti, you own everything.
Given your goal of a quick pilot with clear success metrics, I'd recommend the pure cloud service path (ZScaler or Cloudflare). The control OpenZiti offers isn't worth the overhead for a 10-user proof of concept. My specific pick would depend on your existing vendor relationships: if you already use ZScaler for web gateway or Cloudflare for DNS/CDN, go with that vendor for integration simplicity.
Your question about the agent is the right one. It's the main thing your power users will actually notice and complain about.
If you're testing ZScaler or Cloudflare, demand to run the posture check agent separately first. Don't bake it into the initial auth flow. Let them get used to the new access method, *then* roll the agent out a week later. You'll get cleaner feedback on each piece.
Also, define "lightweight." Check its CPU impact on your devs' machines during builds. That's a real-world gotcha that'll kill adoption faster than any setup complexity.
metrics not myths
That's a solid operational tip. Separating the agent rollout is good pilot hygiene.
To your point about defining "lightweight," I'd add that memory footprint matters just as much as CPU for some of our engineers running IDEs with multiple projects open. Have you seen any data on agent behavior when a machine is already under heavy load? It's one thing to check it on a fresh reboot, another when the system's strained.
Review first, buy later.