Skip to content
Notifications
Clear all

My workflow to clean up stale leads saves 4 hours a week

6 Posts
6 Users
0 Reactions
30 Views
(@jackson)
Estimable Member
Joined: 3 months ago
Posts: 82
Topic starter   [#13936]

Our team manages a large number of Flux Kustomizations across multiple clusters, which naturally accumulate stale objects—resources that exist in-cluster but are no longer present in the source Git repository. These stale "leads," particularly ConfigMaps and Secrets referenced by other resources, were causing reconciliation errors and cluttering our observability dashboards. Manual cleanup was a recurring, time-consuming task.

I developed an automated workflow that identifies and removes these resources, saving our platform team approximately four hours of manual work per week. The core of the solution is a combination of `kubectl` pipelines and a simple script executed by a Kubernetes CronJob.

The process has two phases. First, we generate a list of all Flux-managed resources. Then, we compare it against the actual resources in a given namespace, filtering for those with the `kustomize.toolkit.fluxcd.io/name` label but missing the `kustomize.toolkit.fluxcd.io/namespace` label (a common indicator of a stale object). The following script is deployed as a CronJob in the management cluster.

```bash
#!/bin/bash
set -euo pipefail

NAMESPACE=""
CONTEXT=""

# Export all resource names managed by Flux in the namespace
kubectl --context $CONTEXT get kustomization -n $NAMESPACE -o jsonpath='{.items[*].spec.sourceRef.name}'
| tr ' ' 'n' | sort -u > /tmp/flux-sources.txt

# For each source, get all associated resources
while read -r source; do
kubectl --context $CONTEXT get all,cm,secret -n $NAMESPACE -l kustomize.toolkit.fluxcd.io/name=$source -o name >> /tmp/flux-managed.txt 2>/dev/null || true
done /tmp/stale-resources.txt || true

# Safe deletion with dry-run first (remove --dry-run=server to execute)
if [ -s /tmp/stale-resources.txt ]; then
echo "Found stale resources:"
cat /tmp/stale-resources.txt
# kubectl --context $CONTEXT delete --dry-run=server -n $NAMESPACE $(cat /tmp/stale-resources.txt)
fi
```

The CronJob runs weekly, logging the stale resources it would remove. After a few cycles of verification, we enabled actual deletion. This systematic cleanup eliminated the recurring reconciliation errors and reduced alert noise significantly. The key is the precise label selection logic, which reliably isolates the stale objects without risking managed resources.

—J


—J


   
Quote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

Interesting approach using label presence as a proxy for staleness. I've seen similar patterns in data pipeline orchestration where "orphaned" tables lack lineage metadata.

Have you quantified the reconciliation error reduction since implementing this? I'd be curious if the removal rate follows a predictable decay curve, or if it's more sporadic based on team Git activity.

One caveat from our similar cleanup jobs: watch for resources that might be referenced by non-Flux controllers. We added a dry-run mode that logs potential deletions for a week before enabling actual deletion.



   
ReplyQuote
(@elliotk)
Reputable Member
Joined: 2 months ago
Posts: 323
 

That dry-run mode is such a crucial step. We learned that the hard way with a similar RAG cleanup script that was deleting stale vector embeddings. It turned out a separate analytics process was referencing some of those "orphaned" chunks by ID, causing silent failures.

On the quantification point, our error reduction wasn't linear at all. It was super spikey, correlating almost perfectly with major PR merges or dependency updates that refactored a bunch of resources at once. The weekly cleanup just catches the fallout. I'd bet your pattern is similar if team git activity is in bursts.

Your mention of non-Flux controllers is spot on. We added an explicit exclusion list for resources owned by certain ArgoCD ApplicationSets, which are effectively in a different "layer" of GitOps. Have you run into any other specific controllers that trip this up?



   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

The label check is clever. I've used a similar pattern but found it fragile after Flux upgrades. The exact label keys changed once between minor versions.

You should wrap that `kubectl` pipeline in a `--dry-run=client` flag first, outputting to a log. Run it for a few cycles to audit what it catches. I've seen that missing namespace label on resources that were intentionally pruned by Kustomize but still needed for rollback scenarios.

Also, that script will miss any CRDs installed by Flux that later get removed from Git. Those don't always carry the standard kustomize labels. You might need a second pass for `api-resources` with a `fluxcd.io` owner.


Build once, deploy everywhere


   
ReplyQuote
(@jasonl)
Eminent Member
Joined: 2 months ago
Posts: 23
 

Good call on the label key instability. We ran into that after a Flux v2 upgrade where `kustomize.toolkit.fluxcd.io` became `kustomize.toolkit.fluxcd.io/name`. It broke our selection logic for a day.

The CRD gap is a real blind spot. We actually built a separate reconciliation check for that, querying the `fluxcd.io` finalizer on objects rather than labels. It's more verbose but survives label schema changes.

Did you find the dry-run logs gave a lot of false positives from rollback resources, or was it a manageable volume to review?


Data beats opinions.


   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

That's a great point about Flux upgrades breaking the label checks. I hadn't considered version drift at all 😅.

When you ran the dry-run logs, how many false positives did you typically see? I'm worried that reviewing them might become its own time sink.

The CRD angle is really smart, I'll have to look into that owner reference check.



   
ReplyQuote