Skip to content
Notifications
Clear all

Just built a dashboard that shows GitOps sync status across 50 clusters - code shared

47 Posts
47 Users
0 Reactions
59 Views
(@contractor_consultant_mike)
Reputable Member
Joined: 4 months ago
Posts: 329
 

Nice solution. The 5-minute alert threshold on the gauge is a smart choice - it cuts through the noise without losing real issues.

One thing to watch out for is that filtering by project name can be brittle if naming conventions change or if new system projects are added. We ended up filtering based on an annotation (`monitoring.enabled: "true"`) that teams add to their ArgoCD Application CRDs. It's an extra step for devs, but it's more explicit and avoids accidental visibility gaps.

I'm curious, do you track the age of the `lastSync` timestamp? For us, an app stuck in a "Synced" state but with a very old sync time was an early indicator of a stalled reconciliation.


Integrate or die


   
ReplyQuote
(@integration_jane_new)
Reputable Member
Joined: 7 months ago
Posts: 304
 

You've raised valid points about the limitations of boolean gauges for detecting flapping. A state duration metric would indeed provide more granular visibility into intermediate states like `Progressing`. We actually augment our primary gauge with a separate `sync_duration_seconds` histogram that tracks the time spent in non-Healthy states. This catches the intermittent issues you mentioned, but adds complexity to the alert rules.

The scaling concern for 50 clusters is real. We don't use a simple loop with a single HTTP client. Each cluster target has its own independent scraper goroutine with a dedicated connection pool and a circuit breaker. A slow cluster API only trips its own breaker and marks that target down, it doesn't block the others. The architectural overhead was significant but necessary.

On the system project filter creating a hidden dependency chain, I completely agree it's a trade-off. Our compromise was a secondary, less-frequently-scraped dashboard that includes all projects, annotated with ownership. It's not for alerts, but it's the first place we look when a business app shows drift without an obvious cause in its own project.



   
ReplyQuote
Page 4 / 4