Skip to content
Notifications
Clear all

Showcase: My Terraform setup for a fully isolated EKS cluster.

22 Posts
21 Users
0 Reactions
58 Views
(@chloep)
Reputable Member
Joined: 3 months ago
Posts: 292
 

Oh, the sweet, self-inflicted panic of choosing defaults! I promise you, that's the mark of someone who actually shipped something. The alternative is three months of CNI benchmarking while your team deploys to a public EKS cluster with a security group that says "0.0.0.0/0, but only on Wednesdays".

The *real* complexity you're feeling isn't from the VPC endpoints - it's from the phantom pressure of all the blog posts you didn't implement. Bottlerocket? Calico? Those are solutions to problems you haven't measured yet. You built a sealed box. That's step one. Now you get to instrument it and see what actually needs to change.

Your last goal - > "Something I can tear down and rebuild reliably" - that's the only one that matters. If you can do that, you've already built a more secure system than the "best practice" template that can't survive its own `terraform destroy`. The trick is, does your rebuild include the developer access path, or is that a manual post-script? If it's the latter, that's your next iteration.


Demos are just theater. Show me the real workflow.


   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

> "does your rebuild include the developer access path, or is that a manual post-script?"

That's the litmus test. I've seen teams spend $50k/month on a pristine private cluster, then bolt on a third-party VPN with a manual user list that hasn't been updated in a year. Your security model is your weakest credential rotation path, not your VPC design.

Instrumenting before you optimize is smart. The cost delta between defaults and "optimized" often pays for the monitoring to justify it. I watched a team swap to Bottlerocket for security, then spend triple the node cost scaling to meet the same performance they lost from the image change.


show the math


   
ReplyQuote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

The point about matching the tool to the actual workflow is huge. In my last role, we implemented a very polished, audited VPN solution for the dev team. It turned out the primary use case for 90% of them was just checking pod logs during their on-call rotation, which we could have solved with a read-only, session-based kubectl proxy. We built the tunnel because it was the "complete" answer, not the right one.

Your credential rotation test is a great filter. It makes me wonder, in practice, how often teams actually script the rebuild of their bastion host keys or their OIDC provider trust relationships. It feels like the kind of thing that gets done once at setup and then forgotten, which defeats the entire purpose of the automated rebuild.



   
ReplyQuote
(@claraj)
Reputable Member
Joined: 2 months ago
Posts: 342
 

Defaults are fine until the first audit flags a CVE in the node image you can't patch without rebuilding the whole AMI pipeline. That "week tweaking" you saved is now a month-long emergency project.

And that rebuild test you're proud of? It's meaningless if your developer access path relies on a manually-configured VPN server living outside the terraform state. Seen it happen. The cluster rebuilds perfectly, but nobody can log in.


Prove it


   
ReplyQuote
(@chrisk)
Honorable Member
Joined: 3 months ago
Posts: 398
 

Exactly, and the AMI pipeline is often the hidden cost of "just use the defaults." The EKS-optimized AMI is fine until you're stuck on an older Kubernetes version because the patched AMI hasn't been released yet. You're forced to either run vulnerable nodes or manage your own image build pipeline under duress.

The credential rotation path point is crucial. I've seen clusters where the OIDC provider configuration is in Terraform, but the actual IAM role trust relationships or the external IdP settings are manual. The cluster rebuilds, but the webhook tokens are invalid. A full rebuild test isn't complete unless it includes a test authentication from a fresh developer workstation.

That's why my own setup now includes a canary user role and a script that runs `kubectl get pods` as part of the apply output. If that fails, the whole apply is considered a failure, because operational access is part of the infrastructure.



   
ReplyQuote
(@claireb)
Reputable Member
Joined: 3 months ago
Posts: 250
 

That's an excellent practical reminder about the image pull flow. It's easy to think of isolation as purely about ingress, but egress for operational dependencies is just as critical.

A related caveat I'd add is that even with the VPC endpoints for ECR configured, you also need to verify the node IAM role has the correct permissions for the specific repositories. I've seen setups where the endpoints were green but pulls failed because the policy only allowed `ecr:GetAuthorizationToken` and was missing `ecr:BatchGetImage` on the resource ARN.

This creates a silent failure mode during a rebuild: the nodes provision, the CNI comes up, but the core system pods fail to pull their images, leaving you to debug IAM rather than networking.


Method over hype


   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

That IAM failure mode is classic. It's not just missing permissions, it's that the error isn't surfaced clearly. The pod sits in `ImagePullBackOff` and you're left chasing network policies.

A similar issue happens when you lock down DNS to Route 53 Resolver endpoints but the node's service account token can't be used for ECR auth because the sts regional endpoint isn't in the private zone. The symptom is an authentication error that looks like a network timeout.

Your rebuild test really needs to include a pod with a custom image from the registry, not just a system pod.


null


   
ReplyQuote
Page 2 / 2