Hi everyone. I've been running an ArgoCD HA setup (3 replicas) on a small dev cluster for about six months now. Wanted to share my experience, especially around stability and the real cost.
The stability has been fantastic. We had zero downtime related to ArgoCD itself. The failover between instances during a planned node drain was seamless. The main cost for us is the resource footprint. Running three replicas with requests of 1 CPU and 1Gi memory each adds up. It's worth it for our critical path, but for a simple personal project, it might be overkill.
I'm really happy with it overall. Has anyone else run a similar HA setup? Did you find ways to optimize the resource usage without sacrificing reliability? Thanks for any insights!
Great to hear your HA setup's been so stable! That tracks with my experience, too. The failover really is seamless, isn't it?
On the resource cost, you hit the main trade-off. For a dev cluster, I've seen some teams run with two replicas and lower requests (like 0.5 CPU, 512Mi) during non-work hours via a simple scaling schedule. It's not "full" HA, but it cuts the bill without much risk if you're not deploying overnight.
Have you looked at the memory usage patterns on your replicas? I found one of ours was consistently underutilized, so we adjusted the requests down after some monitoring. Saved a bit.
APIs > promises