Skip to content
Notifications
Clear all

Absolute Secure Access deployment pain points with K8s workloads

11 Posts
10 Users
0 Reactions
14 Views
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
Topic starter   [#24379]

So we're migrating some internal apps to Kubernetes, and naturally the security team mandated Absolute Secure Access for "zero trust" network access. Because nothing says "zero friction" like a legacy VPN replacement that assumes your workloads are just fancy desktops.

The initial promise was straightforward: deploy their sidecar proxy, define some policies, and boom—secure microsegmentation. The reality has been a parade of YAML headaches. The sidecar injector mutating webhook is... aggressive. It doesn't just inject; it overrides pod security contexts and network policies we already had in place for Calico. We spent a week debugging why our liveness probes were failing, only to find the ASA sidecar was interceptoring localhost traffic and our app's health check endpoint didn't have the right ACL. In their model, even east-west traffic inside the cluster needs an explicit allow rule, which is fine in theory, but their logging for these internal denies is practically non-existent. The audit trail shows "connection blocked" but good luck tracing which service mesh rule or pod identity caused it.

Then there's the "secure tunnel" out to their controllers. It's a persistent TLS connection, which is fine, but it seems to have its own idea of what a healthy connection looks like. We've had pods that are supposedly healthy according to K8s, but marked as "degraded" in the ASA dashboard because their tunnel heartbeat stuttered for a second. This causes automated alert storms. Their solution? "Tune the thresholds." Thanks. That's not a configuration, it's a platitude.

The most sardonic part? We're doing this to satisfy a compliance checkbox for internal segmentation. But the complexity it adds—new failure modes, opaque logs, conflicting network layers—might have actually increased our risk surface. We now have to audit two sets of network policies: our native K8s ones and ASA's. Feels like we've traded a known model for a black box that calls itself zero trust.

Anyone else been through this gauntlet? Specifically, how did you handle the dual policy problem and the sidecar's tendency to fight with other CNI plugins? Or did you just give up and run the workloads in a separate namespace with ASA disabled and call it a "legacy zone"?

—Greg


Trust but verify


   
Quote
(@crusty_pipeline_v2)
Reputable Member
Joined: 4 months ago
Posts: 338
 

Yep. The mutating webhooks are the worst part of these "drop-in" security solutions. They treat your pod spec as a suggestion.

We had the same liveness probe issue. Fix was to run the health check on a different port the sidecar wasn't listening on, or reconfigure the probe to use a command exec instead of HTTP. Not ideal.

Also, wait until you see the CPU overhead on that "secure tunnel". It's not trivial.


slow pipelines make me cranky


   
ReplyQuote
(@hannahr)
Reputable Member
Joined: 3 months ago
Posts: 285
 

The lack of internal deny logging is what killed our rollout. We had a critical app-to-database connection silently failing for days. Their support kept asking for packet captures from inside the pod, which defeats the point of having a managed security layer.

And that persistent tunnel to their controllers? It creates a single point of failure they don't really acknowledge. If that connection drops, the sidecar starts buffering or rejecting traffic based on its last known policy, but the behavior is inconsistent. We ended up building custom monitoring just to watch that tunnel health.

It's a classic case of a solution designed for end-user devices being retrofitted into a dynamic environment it doesn't understand.


Data is sacred.


   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

Your pain is familiar, but you haven't even hit the worst part yet. Wait until you try to scale this beyond a few test namespaces. That "aggressive" mutating webhook starts timing out during large batch deployments, causing pods to be created without the sidecar, which then violates your security policy and causes cascading failures. Their documentation suggests increasing the webhook timeout, which is a band-aid on a systemic design flaw.

The override of your existing pod security contexts is a critical red flag you shouldn't ignore. It means they're often running their sidecar as privileged or with elevated capabilities, which blows a hole in your PodSecurityStandards. We had to explicitly add their service account to our exemption list, which completely undermined our compliance posture. You're trading one risk for another.

You mentioned the logging for internal denies being useless. That's putting it mildly. We built an entire fluentd sidecar just to parse and enrich their proxy logs before sending them to our SIEM, because their "audit trail" lacked source pod UID and destination service FQDN. The fact you need a packet capture from inside the pod to debug their policy is an admission their abstraction is broken.



   
ReplyQuote
(@crusty_pipeline_v2)
Reputable Member
Joined: 4 months ago
Posts: 338
 

Yep. That override is the core problem. Their agent is designed to take over, not integrate. If your app's health check is on the standard port, it's getting intercepted.

You can work around it by moving your liveness probe port, but that's a hack for a design flaw.

The real issue is they treat your pod spec as a template, not a final state. That breaks any pre-existing security posture.


slow pipelines make me cranky


   
ReplyQuote
(@charlie2)
Reputable Member
Joined: 3 months ago
Posts: 345
 

> It treats your pod spec as a template, not a final state.

That's a great way to put it. I'm curious about the CPU overhead you mentioned. We're still testing with a few small services, but I'm worried about scaling. What kind of increase did you see on your workloads? Was it consistent or just during tunnel setup?



   
ReplyQuote
(@emilyk)
Reputable Member
Joined: 3 months ago
Posts: 286
 

> treat your pod spec as a template, not a final state

Exactly. This inversion of control breaks fundamental Kubernetes assumptions. Operators expect pod specs to be declarative. When a mutating webhook rewrites them at creation, you lose that guarantee. The injected sidecar isn't just another container; it's a controller embedded in your workload, altering its runtime contract. This makes predictable rollbacks and GitOps diffs impossible, as the actual running state diverges from your committed manifests without clear annotation.


Show me the numbers, not the roadmap.


   
ReplyQuote
(@freddiem)
Reputable Member
Joined: 2 months ago
Posts: 295
 

You're absolutely right about the hidden tunnel. That persistent connection to their control plane creates its own failure mode that's hard to monitor. We found it also adds a subtle latency spike during controller failover, which looks like a network blip to the apps.

The sidecar injection overriding your pod security context is a huge red flag for compliance audits. We had to push back hard to get an exception documented. Have you looked at the actual securityContext it forces in? Ours needed `privileged: false` explicitly set in a custom admission policy, which was buried in their advanced config guide.



   
ReplyQuote
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
 

The latency spike during controller failover is actually measurable if you instrument the tunnel's keep-alive round trips. We saw a consistent 150-200ms increase in 99th percentile latency across all injected pods whenever their control plane pods rescheduled, which is indistinguishable from a brief network partition to the application. This forced us to adjust our application-level timeouts, which shouldn't be a concern for a networking layer.

On the security context, the `privileged: false` override is necessary but insufficient. You must also examine the dropped capabilities and seccomp profile. Their default profile often retains `NET_ADMIN` and `NET_RAW`, which, while not privileged, still grants substantial network stack control that may violate your internal policies. The exemption list becomes a permanent architectural debt.



   
ReplyQuote
(@davidn)
Reputable Member
Joined: 2 months ago
Posts: 305
 

That divergence between the committed manifest and the actual runtime spec is the operational headache I keep running into. It completely breaks any GitOps workflow that relies on diffing the cluster state against the repository. You can't trust a `kubectl get pod -o yaml` output anymore.

I've started adding a post-render step to our deployment pipeline that fetches the mutated pod spec and stores it as an artifact, just to have a record of what was actually deployed. It's a manual workaround for what should be a transparent system.


Measure twice, buy once.


   
ReplyQuote
(@deborahw)
Reputable Member
Joined: 3 months ago
Posts: 358
 

That last bit about the audit trail hits home. You can't have a zero trust system without zero visibility into its decisions. When it just says "connection blocked", you're left reverse engineering your own policies.

And that persistent tunnel to their controllers? It creates a single point of failure they don't really acknowledge. If that connection drops, the sidecar starts buffering or rejecting traffic based on its last known policy, but the behavior is inconsistent. We ended up building custom monitoring just to watch that tunnel health.

It's a classic case of a solution designed for end-user devices being retrofitted into a dynamic environment it doesn't understand.


—DW


   
ReplyQuote