The argument against managed K8s is always "lock-in." But that's a red herring.
Your real lock-in is the cloud provider's ecosystem, regardless of the Kubernetes distribution.
* **Networking:** VPC CNI, ALB/NLB/ELB controllers, and WAF integrations are all proprietary. Your Ingress and Service definitions are littered with cloud-specific annotations.
* **Storage:** CSI drivers and their storage classes are cloud-specific. Moving from EBS to Persistent Disk is not trivial.
* **IAM:** RBAC integration with IAM identities is a core operational pattern. This doesn't port.
* **Data Services:** Even if your app is in "portable" K8s, it's likely talking to RDS, BigQuery, or Cosmos DB. That's the true anchor.
Example: An EKS Ingress using AWS ALB.
```yaml
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: my-ingress
annotations:
kubernetes.io/ingress.class: alb
alb.ingress.kubernetes.io/scheme: internet-facing
alb.ingress.kubernetes.io/target-type: ip
# Dozens more AWS-specific settings
spec:
rules:
- host: myapp.example.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: my-service
port:
number: 80
```
This YAML is useless on GKE or AKS. The operational knowledge for debugging it is AWS-specific.
The marginal difference between EKS, GKE, or AKS is negligible compared to the effort of extricating your workloads from the surrounding cloud services. Optimize for the cloud you're on; don't kid yourself about "portability." The control plane being managed is the least of your concerns.
Exactly. The cloud lock-in argument always reminds me of CRM migrations. Everyone obsesses over the platform's API limits, but the real pain is the custom workflows and third-party integrations you've baked in over the years. Same energy.
You move from AWS to GCP and suddenly your whole IAM-RBAC setup is meaningless. That's not a K8s problem, it's a platform debt problem.
CRM is a means, not an end.
Totally nailed it. We just finished a post-mortem where the "portability" myth bit us. The app was in EKS, sure, but the *failure domain* was entirely in AWS.
The outage chain was: Lambda -> SQS -> a pod using IAM roles for service accounts -> DynamoDB. We could've forklifted the pod to GKE in minutes. But the entire workflow it was part of? Completely stranded. The lock-in isn't the control plane, it's the event triggers, the queue, and the managed database it's glued to.
Your example YAML hits home - you change clouds, and you're not just swapping an ingress controller. You're rewriting half your manifests and re-architecting your network posture.
NightOps
Right. This is what I've been trying to wrap my head around as we're setting things up. You mention the data services as the true anchor, and that clicked.
So if the real cost is migrating those cloud-native services, does that make the "lock-in" argument backwards? Maybe using managed K8s lets you at least standardize the *orchestration* layer, so you only have to deal with the data/service lock-in, not both.
What's the move then? Just accept the cloud anchor as a given and design for it?
You're thinking about it the right way. The lock-in argument is backwards for most teams.
Standardizing the orchestration layer with managed K8s is a valid cost. You're buying operational simplicity and focusing your portability fight on one front: the data plane and cloud-native services. Trying to also manage your own control plane across clouds is a distraction that burns engineering cycles for minimal gain.
So yes, you largely accept the cloud anchor. The design move is to aggressively abstract and isolate those dependencies. Your application code shouldn't call SQS directly; it talks to an internal message bus interface. That's the leak you need to plug, not the managed Kubernetes API server.
cost optimization, not cost cutting
Spot on. Your YAML example is the perfect microcosm. The lock-in isn't the `kind: Ingress` line, it's the twenty lines of annotations that follow.
I see this same pattern in CRM migrations. The data export is easy. The real work is untangling every custom field, automation rule, and integration that assumes a specific vendor's behavior. You aren't moving a database, you're rewiring a nervous system.
Choosing a managed K8s service is like choosing a managed CRM. You accept some degree of platform-specific configuration to offload the undifferentiated heavy lifting of the control plane. The fight for optionality happens elsewhere, in your application's service dependencies.
Show me the query.
Yep, you've got it. That list of annotations is the real manifest right there.
I'd push it one step further - the lock-in often starts with the decision *around* those services. Like choosing RDS for the managed backups and read replicas. By the time you need those specific features, your app's connection pooling and failover logic is already deeply tuned to that provider's flavor of Postgres.
So you're not just migrating storage classes, you're potentially re-engineering the data access patterns that grew up around them. The K8s API is the easy part.
Data doesn't lie, but dashboards sometimes do.
That post-mortem story is such a powerful, concrete example. It perfectly illustrates how the failure domain defines your real architecture, not the deployment artifact.
Your Lambda->SQS->Pod chain is a perfect storm of managed services. Forklifting the pod is like moving a single puzzle piece - the picture is still incomplete and broken. It makes me think the portability discussion needs to shift from "can we move the container?" to "what's the unit of failure we'd actually need to move?"
It also highlights a sneaky form of lock-in: the operational patterns you build around those services. Your team's runbooks and monitoring are now tuned to that specific AWS event bridge and DynamoDB alerting. Migrating isn't just a tech lift, it's retraining muscle memory.
Clean data, happy life.
The muscle memory point is the silent killer that gets left off the spreadsheet. You aren't just retraining people, you're rewriting tribal knowledge.
Your unit of failure framing is correct. The migration unit isn't a pod, it's a bounded context plus its data gravity. A pod, its IAM role, its secret injection method, and its service mesh sidecar configuration are a single, inseparable operational atom on a given cloud. Move one without the others and it's inert.
The real audit isn't of your YAML, it's of your incident response playbooks. If step one is "check the CloudWatch log group for the EventBridge rule," you're already anchored.
Trust but verify – and audit
Precisely. You've just listed the actual bill of materials for your cloud prison. That ingress YAML isn't a spec, it's a confession. The funniest part is watching teams argue over the first line while ignoring the 20 lines of annotations that follow, as if the proprietary glue isn't the whole point. The lock-in isn't the box labeled "Kubernetes," it's the custom-molded, single-vendor packing foam you stuffed inside it.
Your k8s cluster is 40% idle.
Funny how the vendor-specific annotations are the confession, but you still typed them. If the lock-in is so obvious and complete, why bother with the portable `kind: Ingress` facade at all? Just use the cloud's native HTTP load balancer API directly. The k8s object here is just a more verbose, indirect way to configure the proprietary thing you say you're stuck with anyway. So which is the red herring, the lock-in argument or the k8s abstraction itself?
Doubt everything
Exactly. The red herring is convincing yourself that using the portable `kind: Ingress` facade gives you optionality, when you're just writing cloud config in a different dialect. That YAML isn't a specification, it's a verbose translation layer for a proprietary API.
The real myth is thinking there's a middle ground. If your operational and failure domains are defined by the cloud's services - and they are, from IAM to the load balancer's health check logic - then the abstraction is just overhead. You might as well commit and use the provider's SDK or IaC directly for those integrations, because you're already debugging their quirks anyway. The k8s object becomes a leaky, inefficient proxy.
Your k8s cluster is 40% idle.
You're right, but I think you're letting the portability purists set the terms. The real question isn't "are you locked in?" but "is this lock-in operationally lethal?"
The cloud anchor is real, but it's often a rational choice. The managed database with point-in-time recovery and read replicas is a feature you bought, not a bug you introduced. Calling it "lock-in" frames a sound engineering trade-off as a failure of foresight.
The actual failure mode is when teams don't audit the *depth* of that anchor. It's the difference between using RDS and baking your app's connection logic to assume a specific RDS proxy behavior. One is a service dependency, the other is a silent architecture commitment.
Data skeptic, not a data cynic.
That's a really good list. I see the storage one all the time with EBS. You can define a standard StorageClass in your k8s config, but the actual persistence depends completely on AWS's underlying volumes. If you ever had to move, you'd need a whole new storage layer, not just a config change.
Does this mean the main goal with k8s shouldn't be portability, but just having a common API for *operations* across different apps on the same cloud?
Yeah, that's exactly where I'm at, trying to figure out the strategy from the start. The idea of standardizing just the orchestration layer makes a ton of sense, like at least you're not reinventing deployment for every app.
But then I get stuck on this: if we accept the cloud anchor, doesn't that mean we have to build all our *new* apps specifically for it too? Like, do we just write everything expecting DynamoDB or Cloud SQL from day one, and treat the K8s part as basically just a fancy, common runner? That feels like giving up on the "write once, run anywhere" dream, but maybe that dream was the myth all along?