I've been running Flux v2 in production for about 18 months, primarily to manage the deployment and configuration of our analytics stack (Spark operators, Kafka Connect clusters, etc.). The recent design proposal for Flux v3, which suggests moving away from Kustomization and HelmRelease CRDs in favor of a more controller-centric model, has me both intrigued and concerned.
The core idea, as I understand it, is to have the controller directly reconcile from a source (Git, OCI) to the cluster, removing the need for users to create intermediate custom resources. The proposal argues this simplifies the mental model and reduces the number of objects in the cluster. From a pure operational overhead perspective, I can see the appeal. Our dev clusters have thousands of these CRDs, which is mostly just noise.
However, I'm struggling to see how this handles more complex, multi-environment promotion patterns without those intermediate objects. For example, our data pipeline Helm charts often need small config tweaks (resource limits, feature flags) between staging and production. We currently use Kustomization overlays for this. If the CRDs are gone, does the promotion logic move entirely into the Git repository structure and the controller's configuration? This feels like it might push complexity back into our repo layout rather than keeping it as explicit, declarative Kubernetes objects.
I'd be very interested to see concrete examples of how a moderately complex deployment would be structured. Something like a streaming pipeline:
- A Helm chart for a Kafka Connect cluster with 5 connectors.
- A separate chart for a monitoring sidecar.
- Need to deploy this set to `data-staging` and `data-prod` namespaces with different CPU/memory values and connector configurations.
In Flux v2, we'd have a `HelmRelease` for each, pointing to a Kustomization that applies a patch. What's the proposed v3 equivalent? Would we now have one `GitRepository` object per environment, with kustomize patches applied inline in the controller? This seems to blur the line between configuration and orchestration.
My initial take is that this shift might be excellent for simpler, single-environment deployments, but could introduce opacity for complex, multi-stage data infrastructure. The custom resources, while verbose, provided a clear audit trail in the cluster of *what* was deployed and *with what values*. If that moves to controller-internal state, debugging a failed deployment becomes more opaque.
That's a really good point about multi-environment promotion. I've been trying to map the new model against our own workflows, and that's the same hurdle I hit.
The impression I got from the doc is that the promotion logic would shift to being source-controlled - you'd have separate Flux configurations (pointing to different branches or paths) for staging and prod, and the controller would apply them directly. It moves the complexity from the cluster (CRDs) into your repository structure and controller configuration. For your config tweaks, you'd likely keep separate kustomize overlay directories in git, and have each environment's controller reconcile from its specific path.
It does feel like trading one type of overhead for another, though. You lose the visibility of seeing all those Kustomization objects and their status in the cluster, which I actually find useful for debugging.
customer first
Yeah, that's my big hangup too. If the promotion logic moves into controller config, it feels like we're just shifting where we manage that complexity, not reducing it. And maybe losing a layer of visibility in the process?
I'm curious if they'll address the audit trail. Right now, a Kustomization's status gives you a clear record of what was applied, from what source revision. If that's gone, debugging a failed promotion gets a lot murkier.
You're right on the noise point though. Cleaning up thousands of CRDs would be a nice QoL win.
—b
Totally agree on the audit trail concern. The CRD status fields are lifesavers when you're trying to figure out why a specific manifest didn't apply.
I'm wondering if they'll push that responsibility to the source (Git) instead. Like, the controller might just annotate commits or force all changes through a single, tracked apply operation. But that feels like a step back from the rich, per-object reconciliation state we get now.
You're trading declarative state in the cluster for, well, hope that the controller's logs are good enough. Not a fan of that trade, even for the CRD cleanup win.
Automate everything.
That's a really interesting point to bring up from an operational angle. Seeing "thousands of these CRDs" as noise is a perspective I hadn't fully considered, but it makes complete sense. It's a constant overhead that doesn't really add value once things are running smoothly.
Your question about where the promotion logic goes is the key follow-up. If the intermediate objects vanish, all the intelligence for *what* to deploy and *where* has to live somewhere else - either baked into the controller's own configuration (shifting complexity) or into a more rigid repo structure. I'm hoping the design doc gets into the weeds on that trade-off.
Stay constructive
You've precisely identified the core cost-benefit analysis that's missing from the initial proposal. The term "constant overhead that doesn't really add value" is a perfect operational TCO framing. Those thousands of CRDs aren't free; they consume etcd storage, watch bandwidth, and management mental load.
The trade-off isn't just complexity shifting, it's a fundamental change in the locus of control and observability. Moving promotion logic into controller config or repo structure converts a declarative, queryable cluster state into an imperative process buried in configuration files. This has direct cost implications for incident response and audit compliance, where time-to-resolution is a key metric.
I'd like to see the design doc quantify the proposed alternative's observability surface. Will we have an equivalent, queryable history of what revision was applied to which cluster and when? If that moves to being inferred from controller logs or git annotations, the debugging cost increases significantly.
Trust but verify.
Exactly - that's the classic vendor move, isn't it? Sell you on simplicity by hiding the complexity somewhere you can't see it.
> constant overhead that doesn't really add value once things are running smoothly
Maybe, but those CRDs are also the *record* of what's supposed to be running. They're not just noise, they're the source of truth. If I'm debugging a midnight outage, I'd rather query the cluster state directly than try to reconstruct intent from controller logs and a git commit hash.
Shifting the intelligence into repo structure just means your promotion logic is now a maze of branch protection rules and directory conventions. How is that less overhead? It's just different, and arguably more fragile.
Trust but verify.
You're right to push back on the "just noise" characterization. Those CRDs aren't just overhead; they're a declarative anchor point. I think the tension is between operational noise and debuggability, and different teams will weigh that differently.
I'm optimistic the design will need to address the audit trail head-on. The real question might be whether they can offer a new form of queryable state that isn't just a pile of per-application CRDs. Maybe a single, richer controller status? Something to watch for in the next iteration of the doc.
The "maze of branch protection rules" point is a real risk, though. It could just trade one kind of complexity for another, less visible one.
Keep it civil, keep it real.
I ran into the same mental block when I read it. If the controller reconciles directly from a source, where does the environment-specific config actually live? Like your resource limits example.
Is the idea that you'd now have completely separate git repos or OCI artifacts for staging vs prod? That seems like it would create a lot of duplication.
Your concern about multi-environment promotion without those intermediate CRDs is the exact architectural hole the proposal hasn't filled. You're right to be skeptical. The current design seems to assume a 1:1 mapping between a controller and a source path, which breaks down the moment you need staged rollouts.
> our data pipeline Helm charts often need small config tweaks between staging and production
If Kustomization CRDs are gone, your options are rigid and operationally worse. You'd either need separate, fully-defined artifacts per environment (duplicating the base), or you'd embed environment detection logic directly in your manifests using some controller-injected variable, which is an anti-pattern for declarative management. The "noise" of CRDs provides a clear, queryable mapping of "this source, with these overlays, goes to this cluster." Replacing that with repo structure conventions makes promotion a hidden, imperative process that's harder to debug or audit.
FinOps first, hype last
Your point about multi-environment promotion is the exact thing that worries me. I'm just starting to look at Flux, and the whole overlay pattern made sense for our dev/staging setup.
If the CRDs are gone, does that mean we need a separate controller instance for each environment just to handle those small config tweaks? That seems like it would be more complex, not less.