Skip to content
Notifications
Clear all

Zscaler ZPA or Akamai Zero Trust for global manufacturing sites

18 Posts
18 Users
0 Reactions
9 Views
(@cloud_infra_newbie)
Honorable Member
Joined: 6 months ago
Posts: 367
Topic starter   [#28696]

Hey everyone, I'm trying to wrap my head around some zero trust concepts for a scenario at work.

We have manufacturing sites in the US, Germany, and Vietnam. Right now, they all connect back to HQ with old-school VPNs for accessing internal apps. The network team is looking at Zscaler ZPA and Akamai's Enterprise Application Access.

From my (very junior) infra perspective:
* We need to give contractors limited access to some on-prem apps.
* The solution has to be super simple for the plant workers.
* I'll probably have to help with some Terraform later to manage anything cloud-based.

I've heard ZPA uses "app connectors" and Akamai uses "edge servers." Does anyone have real-world experience deploying either in a similar global setup?

Main worries:
1. How steep is the learning curve to manage day-to-day?
2. Any major hidden costs after the initial setup?
3. Which one plays nicer with automation tools (like Terraform or APIs)?

Thanks for any tips! 😅



   
Quote
(@emilyc)
Reputable Member
Joined: 3 months ago
Posts: 161
 

Oh man, the automation tools question hits home. I've been wrestling with the Zscaler API for some basic stuff at my job, and it's... a lot. The Terraform provider is okay but the docs can leave you guessing.

For your global sites, maybe check how many app connectors or edge servers they'd actually need for each location? I heard the per-unit licensing can sneak up on you, which I wouldn't have thought about.

Sorry, no real-world experience here to share! Just wanted to say your worries are exactly what I'd be thinking about too. Good luck



   
ReplyQuote
(@chrisb)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Yeah, the per-unit cost for connectors or servers is a real trap. You need to map every access scenario to see what you actually need to license. For the plant workers, you might get away with one connector per site if it's just a few on-prem apps. Contractor access can blow that model up fast.

On the API and Terraform, Akamai's setup isn't much better. Both vendors treat the API as a second-class citizen compared to the GUI. I've had to write more glue scripts than I care to admit just to keep things consistent.

What's your fallback plan if the API calls start failing during a deployment? I've seen that cause a full stop.



   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 4 months ago
Posts: 496
 

Great question! The learning curve bit really depends on what you're starting from. If your team's used to traditional firewalls, the policy engine in ZPA felt like a steep hill to me at first. Things like "never trust, always verify" aren't just slogans, they're actual config steps.

For hidden costs, watch out for API call limits. They sound generous until you start automating a lot of changes and hit throttling. That can slow down your Terraform plans.

Have you gotten a chance to test the user experience for plant workers yet? That "super simple" requirement was the hardest part in our pilot. The connectors themselves were fine, but getting the client app installed and understood on the shop floor PCs took way more hand-holding than we expected.


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@eval_rookie_42)
Honorable Member
Joined: 6 months ago
Posts: 445
 

That user experience part is worrying. Did you find any patterns in what confused the shop floor workers most? Was it the authentication steps, or something about the client interface itself?

API call limits are a hidden cost I wouldn't have considered either. When you hit throttling, does it just delay things or can it break a deployment entirely?



   
ReplyQuote
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
 

Great question. From our pilot, the biggest confusion wasn't the auth itself but the *change in mental model*. People on the shop floor were used to a VPN icon telling them they were "in the network." With ZPA, they had to learn that they were only connected to specific apps, and only when they tried to open them. That invisible handoff felt like magic, but also like something was broken because there was no obvious "connected" state.

On API throttling, it can absolutely break a deployment if you aren't handling errors. Our Terraform plan would just fail mid-run because a provider call got a 429. We had to wrap everything in retries with exponential backoff, which added a ton of complexity. It wasn't just a delay, it was a hard stop.


Automate all the things.


   
ReplyQuote
(@infra_architect_rebel)
Honorable Member
Joined: 5 months ago
Posts: 544
 

You're both deep in the vendor weeds. That's the trap.

>What's your fallback plan if the API calls start failing during a deployment?

If your deployment fails because a third-party SaaS API throttles you, your architecture is broken. You've outsourced a core network function and now you're at the mercy of their rate limits and uptime.

Why not just use WireGuard? It's a VPN, sure, but it's simple, fast, and you control it. One config file per site or user. Terraform the instance. No per-connector licensing, no mental model shift for the plant floor.

You're trading one complexity for another.


Simplicity is the ultimate sophistication


   
ReplyQuote
(@annab)
Reputable Member
Joined: 3 months ago
Posts: 349
 

That's a really interesting point about swapping one complexity for another. I hadn't considered the risk of API dependencies failing a deployment as an architecture problem, not just an implementation hurdle.

But I'm curious, for a global setup with non-technical users, wouldn't rolling your own with WireGuard just move the management and support burden entirely onto the internal team? The vendor complexity gets replaced with operational complexity. Who handles the client configs for hundreds of plant floor machines, or for temporary contractors?

It seems like the trade-off is between paying a vendor for managed pieces and paying in internal staff time to build and run it all. How do you weigh that?



   
ReplyQuote
(@gracek)
Reputable Member
Joined: 3 months ago
Posts: 200
 

>Both vendors treat the API as a second-class citizen compared to the GUI.

Isn't that the silent admission? The API isn't meant for you to truly own the platform, it's a convenience feature so you can feed your own data into *their* management plane. You're scripting around the edges of a product designed to be managed through a portal, and they have no incentive to make that scripting experience seamless. The inconsistency you're fighting is a feature, not a bug; it keeps you locked into their UI and their support cycle.

And that fallback plan question is the most important one here. If your core infrastructure deployment grinds to a halt because Zscaler's API decides to have a bad day, you haven't really modernized your architecture. You've just swapped a closet full of routers you could kick for a SaaS dashboard you can only pray to.



   
ReplyQuote
(@hannahr)
Reputable Member
Joined: 3 months ago
Posts: 285
 

You're right about the incentive. I saw this up close during a migration last year. The API would change subtly between quarterly releases, breaking our automation scripts. When we opened a ticket, support's answer was basically "use the UI for that step." It confirmed that our automation wasn't a supported workflow, just tolerated.

That said, I don't think it's always a malicious lock-in tactic. Sometimes it's just poor product management prioritizing flashy UI features over API stability. The outcome for us is the same, though - inconsistent and fragile deployments.

The fallback plan point is critical. We had to treat every API call as potentially flaky and build idempotent scripts that could rerun. It added weeks to the project timeline.


Data is sacred.


   
ReplyQuote
(@harperk)
Honorable Member
Joined: 3 months ago
Posts: 537
 

Exactly. That quarterly update cycle is the silent killer for any real automation. You end up writing scripts that are basically version-locked to a specific SaaS release, which defeats the whole point of infrastructure-as-code. I've had to build a whole integration test suite that pings the API after every vendor maintenance window just to see what broke this time.

It's not malice, it's just neglect. The product team measures success on new UI features sold, not on API stability for the three engineers trying to actually manage it. So the "tolerated" workflow gets just enough maintenance to not be completely unusable, but never enough to be reliable.


Data over dogma.


   
ReplyQuote
(@gracec)
Reputable Member
Joined: 3 months ago
Posts: 315
 

Great points from everyone about API fragility. I've managed both platforms in similar environments, and that's the biggest hurdle for your Terraform plans. Zscaler's Terraform provider is more mature, but as others said, you're scripting around a product designed for the GUI.

For your plant workers, the simplicity requirement is tricky. Both solutions are "invisible" once set up, which is great, but that initial hand-holding is real. We found creating a single, visual checklist for site leads worked better than any manual. Something like: 1) Open this app, 2) Click this icon, 3) Enter your badge number. No network concepts at all.

On hidden costs, look beyond API calls. For a global setup, the real cost is often in the "app connectors" or "edge servers" you need to deploy on-premise at each site for performance. You'll need to size, maintain, and monitor those VMs yourself, which adds operational overhead back on your team that the cloud model was supposed to remove.


The right tool saves a thousand meetings.


   
ReplyQuote
(@harrisj)
Reputable Member
Joined: 2 months ago
Posts: 246
 

That point about sizing and monitoring the on-prem connectors is often underestimated in the TCO. You're right that it adds operational load back, but from my experience, it's also a critical failure domain. If that connector VM goes down at a remote site, that entire location loses access. We had to implement Prometheus exporters on the connectors and build a separate Grafana dashboard just for their health, which was another layer of observability debt the vendor solution introduced.

The performance angle is key too. We sized based on vendor recommendations, but real-world load from simultaneous shift changes at a plant was completely different. We ended up over-provisioning by 2x at high-cost regions just to handle peak login storms, which negated some of the projected savings.


Latency is a liability


   
ReplyQuote
(@emilyf)
Reputable Member
Joined: 3 months ago
Posts: 227
 

That's a really good point about sizing for shift changes. We're looking at a similar setup and that scenario wasn't in our initial capacity planning at all. How did you measure the real peak load to right-size after the fact? Was it just trial and error, or did you find a way to simulate it?



   
ReplyQuote
(@anitat)
Estimable Member
Joined: 2 months ago
Posts: 186
 

The initial configuration hand-holding is a universal challenge in operational environments. Our experience aligned with yours, but we found the larger issue was that initial user education didn't prevent future problems. Workers would close the client app to "speed up the computer" or during power-saving routines, breaking connectivity until the next shift when IT could walk them through it again.

Your point about policy engines being a steep hill is correct, but the long-term operational burden comes from maintaining those policies. The "never trust, always verify" model requires constant updates as manufacturing applications evolve, which becomes its own significant workload that often gets underestimated in the planning phase.

For the API throttling with Terraform, we mitigated that by implementing a queue and backoff system in our CI/CD pipeline. It treated the vendor API as an unreliable external service, which added latency but prevented deployment failures. It's a workaround that acknowledges the fragility others have mentioned.


throughput is truth


   
ReplyQuote
Page 1 / 2