Hey folks! Just wanted to share something I've been tinkering with that solved a real headache for my team. We run a bunch of Rancher-managed K8s clusters across different environments (dev, staging, prod), and we kept running into config drift between projects. Manually syncing things like ConfigMaps, Secrets, and even some Annotations was becoming a full-time job.
I built a simple Kubernetes operator to automate syncing specific resources between a source project (like our "golden" staging) and target projects. It's essentially a watcher that reacts to changes and replicates them, with some basic filters. Here's the core of the reconciler logic:
```go
func (r *ConfigSyncReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
var syncRule configSyncv1alpha1.ProjectSyncRule
if err := r.Get(ctx, req.NamespacedName, &syncRule); err != nil {
return ctrl.Result{}, client.IgnoreNotFound(err)
}
// Fetch source resource
var sourceResource unstructured.Unstructured
sourceResource.SetGroupVersionKind(syncRule.Spec.SourceGVK)
err := r.Get(ctx, client.ObjectKey{
Namespace: syncRule.Spec.SourceNamespace,
Name: syncRule.Spec.SourceName,
}, &sourceResource)
if err != nil {
// Handle error or wait for source to exist
return ctrl.Result{RequeueAfter: 30 * time.Second}, nil
}
// For each target namespace in the rule...
for _, targetNS := range syncRule.Spec.TargetNamespaces {
// Create or update the resource in the target
targetResource := sourceResource.DeepCopy()
targetResource.SetNamespace(targetNS)
// Strip unwanted fields like UID, ResourceVersion
cleanForSync(targetResource)
if err := r.Ensure(ctx, targetResource); err != nil {
r.Log.Error(err, "failed to sync to namespace", "namespace", targetNS)
}
}
return ctrl.Result{}, nil
}
```
The custom resource definition looks like this:
```yaml
apiVersion: configsync.example.com/v1alpha1
kind: ProjectSyncRule
metadata:
name: sync-critical-config
namespace: rancher-project-staging
spec:
sourceGVK:
group: ""
version: v1
kind: ConfigMap
sourceName: app-config
sourceNamespace: staging-app
targetNamespaces:
- prod-app
- qa-app
```
Why did we go this route instead of other tools?
* **Helm/ArgoCD:** Felt too heavy for just config snippets; we wanted near-instant sync.
* **Rancher's native tools:** Didn't give us the fine-grained, resource-specific control we needed across projects.
* **Manual `kubectl` scripts:** Error-prone and never run on time.
It's been running for a few weeks now and has already saved us from a couple of "why is prod different?" incidents. Has anyone else tackled this problem? Curious if you used a different approach or if you see any pitfalls I might have missed.
Dashboards or it didn't happen.
Oh, this is super timely! We've been battling the same config drift issue, especially with secrets across dev and staging. Doing it manually is a total time sink.
I'm really curious, how are you handling conflict resolution? Like, what happens if someone makes a change in the target project after your operator syncs? Does it just get overwritten? That's the part that always makes me nervous.
Thanks for sharing the logic snippet, by the way. It's helpful to see the approach.
Conflict resolution is the weakest link in a lot of these sync tools. If you just blindly overwrite, you're asking for a production incident when someone's hotfix gets stomped.
In our implementation, we added an annotation-based lock system and a grace period. Changes made directly to the target resource get a `config-sync/hand-edited: "true"` annotation. The operator sees that flag and enters a reconciliation hold for a configurable duration (we default to 1 hour). It also fires off an alert to the team's Slack channel with a diff of what's being blocked. This gives a human a chance to either revert the manual change if it was a mistake, or to explicitly approve the overwrite by removing the annotation. If the annotation remains after the hold period, the sync fails permanently and requires manual intervention.
It's not perfect, it adds operational overhead, but it prevents silent overwrites. The alternative, trying to merge changes automatically, is a configuration nightmare you don't want.
That annotation-based lock system sounds like a smart safety measure. I've seen similar approaches in CI/CD pipelines where manual deployments block auto-promotions.
I'm curious about the operational overhead you mentioned. How often does that Slack alert actually fire in practice? In our case, manual edits to production configs are rare but high-stakes, so I wonder if the alert fatigue would outweigh the risk of a silent overwrite.
That's a good point about alert fatigue. We have a similar setup for monitoring manual deployments, and the alerts only fire maybe once or twice a month. But when they do, it's usually because something is wrong and needs immediate attention, so people take them seriously.
Maybe the key is making the alert actionable and clear. Ours includes a link to diff the configs and a one-click "Approve Overwrite" button in the Slack message itself, which reduces the overhead.
I like the idea of a grace period, but have you considered making it environment-specific? Like a 5-minute window for dev, but a full 24 hours for prod? That could cut down on noise for the high-stakes environments where changes are truly rare.
This addresses the config drift problem efficiently, but I'm immediately drawn to the operational cost angle. A custom operator introduces significant long-term maintenance overhead: you're now responsible for its lifecycle, security updates, and monitoring.
The cost-benefit analysis depends heavily on the scale of drift you're preventing. If this automates hours of manual toil each week, it's a clear win. However, if the sync rules are few and relatively static, a simpler GitOps approach with a single source of truth in a repository might have been a more maintainable, vendor-neutral solution. Did you evaluate that route first?
That's a really neat approach! I'm just starting to explore Rancher and config management, so this is super helpful to see.
The part about manually syncing ConfigMaps and Secrets being a full-time job really hits home. We're a smaller team and that's exactly the kind of manual work we're trying to avoid.
How did you decide which resources to sync? Are you filtering by label, or is it based on the resource type?
> How did you decide which resources to sync? Are you filtering by label, or is it based on the resource type?
Great question, I was wondering the same thing! Being new to this, I'm also curious about the starting point. Did you just sync everything at first and then pare it back after seeing what caused issues? I feel like knowing what *not* to sync is as important as knowing what to sync.
I'm assuming they used labels, but I'm not sure how you'd handle a resource that needs to sync *sometimes* but not others. Maybe a special annotation?
rookie
Yeah, I had the same thought about what to exclude. I'd guess starting with everything would be noisy and risky, like syncing a secret that's meant to be local to one project.
> Maybe a special annotation?
That's a clever idea. Maybe an annotation like `sync/ignore: "true"` could be the off switch. I'm also wondering if there's a way to define a safe list of resource names or label selectors in the operator's config, so you only sync what you explicitly say.
Filtering by resource type alone is insufficient; you need a compound strategy. Our approach uses a mandatory inclusion label *and* allows explicit opt-outs.
First, we define a base selector like `sync-managed=true`. Any ConfigMap or Secret without that label is ignored. This solves the "what to include" problem.
Second, for the "what to exclude" scenario, we respect an annotation like `sync-managed/exclude: "true"`. This handles the edge case where a resource is labeled for sync but has a temporary, local override. The logic, in pseudocode, looks like this:
```
if resource.labels["sync-managed"] != "true":
ignore
elif resource.annotations["sync-managed/exclude"] == "true":
ignore
else:
sync
```
This gives you a clean audit trail. You can query for all labeled resources to see your sync surface, while the annotation acts as a clear, immediate override without altering the broader selector.
Show me the numbers, not the roadmap.
That's a fascinating solution, and I appreciate you sharing the code snippet. It makes the concept much more concrete.
I have a question about the source project selection. You mention using a "golden" staging environment as the source. Was there a particular reason you chose a staging environment over, say, a dedicated configuration project that doesn't run workloads? I'm thinking about the risk of accidental changes in a live environment propagating out.
Also, how did you handle the initial bootstrap? Did you have to manually apply the sync rules to all target projects, or did you find a way to automate that distribution as well?
Great question about the source project! Using staging as the "golden" source was actually a pragmatic decision for us. It's the last environment before prod where full integration tests run, so any config there has already been validated in a live-but-not-customer-facing cluster. A dedicated config-only project feels a bit too detached from reality for our workflow - we'd risk syncing something that breaks because it was never tested with real workloads.
For the bootstrap, you've hit on a real challenge! We did automate it. The operator itself is deployed via a Helm chart, and the chart values include the source project selector and target label. We use ArgoCD to deploy that chart to all our Rancher-managed clusters, with cluster-specific values setting their own target project IDs. So the rollout was a single PR to our GitOps repo, and Argo handled the rest. It meant the sync rules were deployed alongside the operator itself, which was pretty clean.
— francesc
The conflict resolution is the weakest part. It just overwrites. Always.
If you can't trust your team not to modify the target directly, you've got a process problem no operator can fix. The sync is meant to enforce a single source of truth. Any manual change in the target is, by definition, drift.
read the fine print
Environment-specific grace periods are a solid idea in theory, but they add operational complexity you'll feel in the runbooks. Now you're managing different timeouts per env, and someone has to remember which is which during a 3 a.m. pager alert.
The more actionable your alert, the less you need a variable window. A clear diff and an approve button get the job done. Adding a tiered grace period just means you'll have to explain the matrix to the new hire during the next Sev2.
shift left or go home
That's a clever way to handle bootstrap with Helm and ArgoCD. It sounds like your staging setup acts as both a test environment and a config source of truth, which makes sense.
I'm curious about versioning though. How do you track what config version is currently synced to each target project? Do you tag the source resources somehow?