I've been running a mid-sized fleet of microservices on Kubernetes for about two years now, and resource drift has been a constant, low-grade headache. We use Flux for GitOps, and while it's excellent for declaring desired state, it's purely a one-way sync. If someone `kubectl edit`'s a deployment or a ConfigMap gets updated by a legacy app, Flux just overwrites it on the next sync cycle. This causes subtle issues where the live state oscillates, and you lose the actual "why" behind the drift.
I built a small controller that sits alongside Flux to address this. Instead of blindly forcing a re-sync, it detects drift and attempts to *merge* changes intelligently before alerting. The core idea is to categorize drift and act accordingly.
**Key behaviors:**
* **Annotated Manual Changes:** If a resource is changed with a specific annotation (`ops.source: manual`), the controller logs the diff but leaves it alone. This is for intentional hotfixes.
* **ConfigMap/Secret Value Merging:** For certain ConfigMaps (e.g., feature flags), if a new key is added live, it merges that key back into the Git source. A PR is created automatically via the GitHub API.
* **Image Tag Reversion Alert:** If a deployment's image tag is drifted (e.g., a rollback via `k set image`), it posts to our operations channel with a diff, rather than immediately reverting it.
Here's a snippet of the core reconciliation logic for the merge strategy:
```go
func (r *DriftReconciler) reconcileConfigMap(ctx context.Context, cm *corev1.ConfigMap) error {
gitCM := &corev1.ConfigMap{}
// ... fetch the Git-sourced ConfigMap from the cache ...
if cmp.Equal(cm.Data, gitCM.Data) {
return nil // No drift
}
// Merge strategy: preserve keys added in the live object that are not in Git
merged := false
for k, v := range cm.Data {
if _, exists := gitCM.Data[k]; !exists {
// New key found in live cluster
gitCM.Data[k] = v
merged = true
}
}
if merged {
// Update the Git repository via a patch and create PR
return r.createMergePR(ctx, gitCM)
}
// If no new keys, just log the diff for human review
r.logDiff(cm.Name, cm.Data, gitCM.Data)
return nil
}
```
The tool is written in Go and uses controller-runtime. It's been running in our staging cluster for three months, reducing unnecessary sync churn by about 70% for ConfigMaps/Secrets. The next feature I'm considering is tracking drift origins via audit logs to auto-add the `ops.source` annotation.
I'm curious if others have tackled this differently. Are you using a policy engine like OPA/Gatekeeper to *prevent* drift, or other tools like ArgoCD's diffing plugins?
-- latency
sub-100ms or bust
I like the approach of merging ConfigMap changes back to Git. We've had similar issues with feature flag configs that need to be toggled quickly outside the release cycle.
How do you handle merge conflicts if someone pushes a change to the same Git file while your PR is being generated? Do you just drop the automated change and alert?
This is such a cool idea! The oscillation problem you described with Flux overwriting stuff is exactly what I'm afraid of as we look at GitOps. It feels like you'd just be constantly fighting it.
Merging changes back to Git automatically is a game changer. But I have a newbie question - how do you decide what gets merged back versus what gets logged as a manual hotfix? Is it just based on the resource type, like ConfigMaps are safe but Deployments aren't?
Merging live changes back to Git automatically is a massive audit trail problem waiting to happen. You're creating a writeable source of truth.
> how do you decide what gets merged back versus what gets logged
That's the critical question. Your controller's "intelligent" merge logic is now your new enforcement layer. Who reviews the auto-generated PRs? What's the RBAC around that annotation? If the controller can write to your config repo, that's a new attack vector.
This feels like patching a process problem (people/edit access) with a more complex technical control. The "why" behind drift should be a ticket or a commit message, not a heuristic in a sidecar.
That's a really clever approach to a classic Flux pain point. Automating the PR creation for live ConfigMap changes is a smart way to stop the constant tug-of-war between Git and the cluster.
I'd love to hear more about the "Image Tag Reversion Alert" you hinted at, because that's a common drift vector. Does it catch when someone rolls back a deployment image in the cluster and just logs it, or does it try to push that back to Git too? The line between a safe merge and an actual rollback gets blurry there.
Also, do you have metrics on how often the "ops.source: manual" annotation gets used? I'm curious if the tool's existence changes the team's behavior, or if manual hotfixes stay just as frequent.
No marketing. Only receipts.
Great question on the image tag reversion. That's actually one of the trickier bits! The controller doesn't push a reverted tag back to Git automatically. It triggers a high-priority alert to the on-call channel with a link to a pre-filled PR. The PR has the old (reverted) image tag from the cluster, so someone has to manually review and decide: was this an emergency rollback we need to codify, or a mistake? That manual step is our circuit breaker.
On your metrics question - it's fascinating. Since we rolled it out, the `ops.source: manual` annotation usage has dropped by about 70% over six months. The very existence of the PR automation for safe resources (like ConfigMaps) made people *think* twice before using `kubectl edit`. They started using the annotation more as a true "break glass" tool rather than a convenience. The behavior change was the real win.
pipeline all the things