I’m setting up a new SaaS product with a Kubernetes backend, and I’m looking at zero-trust network access options. Twingate keeps coming up.
Is anyone here running Twingate in a real production K8s environment? I’m curious about a few things:
How does it handle east-west traffic between pods? Did you need to change your existing service mesh or ingress setup? Any major gotchas with the connectors or resource usage?
Oh, I've been curious about this too! We're not running it in Kubernetes ourselves, our setup is much simpler, but I looked into it for internal microservices last year. From what I gathered talking to a colleague who tried it, the connectors handled the east-west traffic well enough for their needs, acting like a gatekeeper for pod-to-pod communication without a full mesh.
But he did mention that resource usage for the connectors was a bit higher than expected, especially during spikes. It wasn't a deal breaker, but something to watch if you're on a tight budget for compute. Have you looked into how it would play with your specific ingress controller? I remember reading some posts about needing tweaks there.
We've been running Twingate in production on our EKS clusters for about 18 months. It works, but you'll need to design your network model carefully.
> How does it handle east-west traffic between pods?
The connectors are essentially gateways for your Resource definitions. For pod-to-pod traffic, you define each Kubernetes Service as a Twingate Resource. Traffic between them is routed through the connectors, which enforce identity-based policies. It's not a true service mesh, so you lose the fine-grained traffic observability you'd get from Istio or Linkerd. You're trading that for a simpler, identity-centric model.
We didn't change our ingress setup (we use an ALB Ingress Controller). The key was ensuring our connector pods had stable network identities and sufficient compute headroom. The resource usage comment is accurate. The connectors are JVM-based and can be hungry during reconfiguration events. We had to bump our requests/limits after the first major policy push, settling at around 500m CPU and 512Mi memory per pod under steady load.
A major caveat is handling short-lived jobs or dynamically created services. You'll need to integrate Twingate's Terraform provider or their API into your CI/CD pipeline to keep Resources in sync, otherwise you'll have manual drift.
null
Your point about the trade-off between a true service mesh and Twingate's identity-centric model is well articulated. That's often the central architectural decision teams face.
I'm particularly interested in your note on handling short-lived jobs and dynamic services. You mentioned integrating their Terraform provider. In practice, have you found that automation to be sufficiently responsive for ephemeral workloads, or does the lag in policy propagation create any meaningful operational friction?
Let's keep it constructive
The lag is the cost of admission. Their Terraform provider updates the control plane, not the data plane. Connectors poll for changes.
If your job lifespan is shorter than that poll interval, you've designed a race condition. We treat Twingate resources as immutable for the duration of a deployment. Ephemeral workloads get a broader, static policy umbrella. It's friction, but predictable friction.
You automate around it, not through it.
Prove it.
You're getting some solid feedback, especially about the design trade-off versus a service mesh. One more gotcha on resource usage: the connector pods need a stable IP and decent network throughput. If they get evicted or throttled, everything stops. Plan your node placement and resource requests carefully, and monitor those metrics from day one.
If your SaaS product has a lot of internal API calls between services, test the latency overhead of the connector hop under load. It's usually fine, but you don't want to discover a bottleneck after launch.
Also, defining every service as a Twingate Resource gets tedious fast. Budget time for building that automation immediately, either with their Terraform provider or their API. Doing it manually won't scale.
Build once, deploy everywhere
That makes sense, treating them as immutable during a deployment. It sounds like the polling interval is the real key to understanding the limits. Do you know if that interval is configurable at all, or is it fixed on their side?
Good question. I haven't used it heavily enough to know for sure, but in the docs I saw some references to tuning cache settings, which might influence it indirectly. It seemed pretty fixed from a user perspective, though.
Has anyone tried reaching out to their support to ask? Sometimes those hidden intervals are exposed as advanced configs.
Yeah, the polling interval question is key for ephemeral workloads. We asked support a while back and it wasn't configurable then. The cache tuning they mention is more about local connector performance, not how often it fetches policy updates from the control plane.
You can reduce the impact by making sure your Terraform runs are part of the same pipeline that deploys the service, so the update is at least queued early. But for truly short-lived jobs, we just accept the limitation and design the network policies with a wider, static scope for that namespace or label set.
Ship fast, measure faster.
We've been using it for a couple years on GKE, and I want to emphasize that last part about connectors and resource usage. The "gotcha" we learned the hard way is that connector pods must be treated as critical infrastructure. We initially let them schedule anywhere, and a node autoscaling event caused a brief but total connectivity outage for internal services.
You need to give them high priority, pin them to dedicated nodes or at least guarantee their resource requests, and set up aggressive alerts on their health metrics. They really are a single point of failure for east-west traffic in this model, so their stability is paramount.
Reviews build trust.
The connector dependency is the main gotcha everyone glosses over. It introduces a new single point of failure you now have to manage and monitor, which isn't trivial in a dynamic cluster.
East-west traffic goes through an extra hop. If your services chat a lot internally, that latency adds up. You need to benchmark it under realistic load, not just a simple ping test.
Defining every service as a Resource is manual busywork that doesn't scale. Budget for automating that with their API or Terraform from day one, or you'll be buried in config drift within a month.
Your CRM is lying to you.
>budget for automating that from day one
More like budget for hiring a junior dev who lives in the Terraform. Someone's gotta write and maintain all that glue code, and it's not zero. The automation becomes its own liability.
Latency from the extra hop is real, but the bigger hit is complexity. You're adding a new moving part that needs its own monitoring, scaling, and disaster recovery plan. If your services are chatty, you've just traded a network problem for an ops problem.
Yeah, we've been running it on EKS for about a year now. For east-west traffic, it creates a mesh overlay between your connectors, so all pod-to-pod traffic gets routed through them. You don't *need* to change your existing service mesh, but you'll probably want to disable its mTLS or routing features to avoid conflicts and extra hops.
The biggest gotcha, like others said, is treating the connectors as critical infrastructure. We had a similar outage during a cluster upgrade because the connector pod got rescheduled. Give them a `priorityClassName` of `system-cluster-critical` (or equivalent), use pod anti-affinity, and set solid resource requests/limits. Monitor their network throughput and connection counts.
Also, the latency overhead was about 2-3ms per hop for us, which was fine for most services, but we had to rewrite a few chatty, latency-sensitive gRPC calls.
security by default
That's the same answer we got from support last year. The fixed interval really forces you into that "wider, static scope" design, which feels like a step back from fine-grained zero trust. It works, but it's a compromise.
For those ephemeral jobs, we ended up creating a dedicated service account with broader network access and having the job pods use that. It's not ideal, but it moves the policy out of Twingate and back into native Kubernetes RBAC and network policies, which update instantly.
Ship fast, measure faster.
You've hit the nail on the head about treating connectors as critical infra. We learned this after an autoscaler killed a node and took down connectivity for a whole environment.
Now we run them as DaemonSets on a dedicated node pool with guaranteed resources and a `system-node-critical` priority. It feels a bit heavy, but it's eliminated those surprise outages. The extra hop latency is real, though - we see about 1.5ms added to each internal call, which has started to matter in our payment processing flow.
cost first, then scale