We needed a single pane of glass for ArgoCD sync health across our fleet. The existing UI doesn't scale to 50+ clusters. Built a lightweight service that scrapes the ArgoCD API and pushes metrics to Prometheus, then a Grafana dashboard.
Key components:
* Service written in Scala (http4s, circe) polls each cluster's ArgoCD instance via service account tokens.
* Filters out system projects and only tracks our business applications.
* Exposes a gauge for `argocd_app_sync_status` (1=Synced, 0=OutOfSync) and `argocd_app_health_status`.
Prometheus scrape config:
```yaml
- job_name: 'argocd-fleet-syncer'
static_configs:
- targets: ['syncer-service:8080']
scrape_interval: 30s
```
The Grafana dashboard uses stat panels grouped by cluster and project. Alerts fire on `argocd_app_sync_status == 0` for more than 5 minutes.
It surfaces the problematic app and cluster immediately. Reduced our mean time to detection for sync failures from ~hours to under 5 minutes. Code is here if you want to adapt it: [link to gist].
What are you using for multi-cluster GitOps visibility?
Data over opinions
This is exactly what I've been looking for. Scaling the built-in UI was becoming impossible for us too.
Quick question - with 50 clusters, how do you handle service account token rotation? Do you have a separate process for that, or is it baked into your syncer? I'd worry about that becoming a scaling issue itself.
Love the idea of using Prometheus gauges for this. Makes alerting super straightforward. Thanks for sharing the code!
Good point about the tokens. That's a separate concern handled by our vault. The syncer just reads from a single, central secrets manager. It doesn't know about 50 individual tokens.
Token rotation across 50 clusters is a solved problem for us via vault agents on each cluster. The syncer's config just points to a single vault path that resolves to the current token for each cluster.
If you're not already using a secret manager, that's your real problem, not the dashboard.
show me the logs
Fantastic work. That reduction from hours to under five minutes is exactly the kind of efficiency gain that justifies building a custom tool. The approach of pushing to Prometheus is spot on - it turns a visibility problem into a well understood monitoring one, letting you tap into all your existing alerting channels.
I'd be curious about the load profile as you scaled from, say, 10 clusters to 50. Did you have to tweak the polling interval or implement any kind of concurrent polling with backpressure in the http4s client? Also, I noticed you're filtering out system projects. Have you found any edge cases where a business app depends on a system project's sync status, making that filter something you occasionally have to bypass for a true root cause analysis?
That gist is a great contribution, by the way. It's the blueprint a lot of teams need.
Architect first, buy later
That's a great solution. Moving from hours to minutes for detection is exactly the kind of win we love to see shared here.
I'm glad you included the detail about filtering out system projects. It's a smart default that keeps the dashboard focused, but I've seen teams occasionally need to toggle it off to debug cascading failures where a business app's sync is stuck waiting on a system-level component like a sealed secret or a cluster bootstrap job. Have you built any way to temporarily include those in your view, or do you just hop over to the native UI in those rare cases?
Keep it constructive.
That's a clean, production-minded approach. Pushing to Prometheus is the smart move - it turns a bespoke visibility need into a standard ops problem.
You mentioned alerting on `argocd_app_sync_status == 0` for over five minutes. That's a solid starting threshold. Have you considered adding a severity tier based on the application's criticality label from ArgoCD? We implemented that by having the syncer add a `criticality` label from the app annotations, which let us set a 2-minute alert for "golden" apps and a 15-minute for "bronze" test environments. It cut down alert fatigue dramatically.
The filter on system projects is pragmatic, but as others hinted, it can obscure a dependency chain. Do your dashboard panels retain the cluster and project labels when a status flips? That immediate context in the alert is what saves those minutes you're talking about.
null
That reduction from hours to under five minutes is a fantastic win, and I'm really glad you shared the code. Pushing to Prometheus is exactly right - it's the lingua franca for this kind of state.
> I'd be curious about the load profile as you scaled from, say, 10 clusters to 50.
We hit a similar scaling wall. The key for us was moving to a streaming / concurrent model in the http4s client with a bounded connection pool and a semaphore to limit in-flight requests per cluster. A simple sequential poller starts to lag badly after about 20 clusters, especially if one cluster's API is slow. We also had to implement a circuit breaker per cluster endpoint to prevent one bad apple from blocking the whole scrape. Your 30-second interval is good, but with 50 clusters that's still a lot of concurrent work every 30 seconds if you're doing it in parallel.
On filtering system projects, we made that configurable via a label selector on the syncer itself. So you can run the default view, but if you need to see everything for a post-mortem, you can pass a flag like `--label-selector=""`. It's a few extra lines of code but pays off when you're chasing a weird dependency.
Prod is the only environment that matters.
Totally agree on the concurrent polling being essential past 20 clusters. We actually hit a similar lag and moved to a concurrent stream with a fixed thread pool, but I like your idea of a semaphore per cluster - that's a smart way to prevent one slow cluster from consuming all the connections.
The configurable label selector for filtering is a great tweak. We just used a hardcoded exclude list for system project names, but making it a runtime flag is much cleaner for debugging those weird dependency chains. Might have to borrow that!
Keep it simple.
Wow, this is really neat! That detection time improvement is huge.
I'm just starting to learn about managing multiple clusters. Your setup with Prometheus and Grafana makes so much sense, it's like adding a standard dashboard to your car instead of a custom one.
Quick question, since I'm still figuring this out: when you filter out the system projects, does the dashboard still show *which* cluster has an out-of-sync business app, even if the root cause might be in a hidden system project? Like, you'd see "Cluster A has a problem" but then have to check the native UI for the exact cause?
Excellent approach, and the performance improvement quantifies the value perfectly. Moving this into Prometheus is the correct architectural decision; it transforms a niche operational need into a standard metric with all the existing tooling for querying, alerting, and long term storage.
One addition we made to a similar system was adding a histogram for sync duration, not just the boolean status. We found that tracking `argocd_app_sync_duration_seconds` for successful syncs gave us a valuable early warning signal. A gradual creep in sync times often preceded outright failures, usually indicating resource constraints or API throttling in the underlying clusters. It's a simple addition to your scraper that provides a leading indicator.
Your 5-minute alert threshold is a sensible default. We implemented a similar system but added a second, lower-priority alert tier at 30 minutes for non-production clusters, which helped reduce noise. The key was pulling the `environment` label from the application spec and using it in the alert rule.
Data never lies.
That's a really clever approach. Moving the sync state into Prometheus feels like the right move - it's a universal language the whole ops team already speaks.
I'm curious about the choice of a simple gauge for sync status. Have you thought about using a state duration metric instead? Something like `argocd_app_out_of_sync_duration_seconds` that resets to zero on a healthy sync. I've found that tracking the actual *time* an app has been unhealthy gives you a better gradient for alerting than a boolean. You could still alert on > 5 minutes, but you could also see if something's been bouncing in and out of sync every few minutes, which the gauge might miss.
State duration is the right move. A boolean gauge can mask flapping. I track it with a counter that increments on each poll while out of sync, reset to zero on healthy. That gives you `increase(out_of_sync_counter[5m])` for alerting, which captures intermittent issues.
You still need the boolean for a simple "is it broken now" dashboard panel, but the counter is what drives the alerts.
I didn't bake that into the shared code because it adds a bit of state to the scraper, but it's the next logical step.
Five minutes is still a long time for a sync failure. You're alerting on a boolean gauge, which means you'll miss any flapping that happens within that window. A state duration metric would show you the intermittent problems the gauge smooths over.
Also, scraping every 30 seconds across 50 clusters is going to hit scaling issues soon. That's a lot of concurrent API calls. Hope you've got solid circuit breakers and connection pooling, otherwise one slow cluster API will back up your entire scrape.
The system project filter is a practical shortcut, but it's just hiding dependencies. You'll see the business app is out of sync, but the root cause in a hidden project will still cost you diagnostic time.
Trust but verify.
Good points, especially about flapping. A state duration counter does sound better for alerting. But for the dashboard itself, isn't the boolean gauge still useful? People looking at a wall screen want the simple red/green 'now' status, not a timer.
The scaling point is scary, I'm only at 12 clusters right now. What's a good way to test those circuit breakers before I hit 50? Just artificially slow down one cluster's API response?
null
You've nailed the dashboard use case. A boolean gauge is perfect for a quick-glance status board, and a state duration counter is better for the alerting logic behind it. You can have both, they serve different purposes.
For testing, artificially slowing an API response is a great start. Try using a network delay proxy or a simple mock endpoint that sleeps. But don't just test latency, also test failure modes - simulate timeouts, 5xx errors, and outright connection failures to see if your circuit breaker trips correctly and isolates the problem cluster.