That's a rough situation, especially during a test phase. Everyone's pointing you to caches and connections, and those are valid, but have you mapped the memory growth timeline against your application deployment schedule itself?
I ask because with 200 routes across 50 nodes, you might have a higher churn of rolling updates or scaling events than you realize. Each endpoint change can trigger a router reload. If those reloads are happening in quick succession, the garbage collector might not get a clean shot at the old config blobs before the next one is built, leading to a staircase pattern of memory use that only goes up.
Before tweaking another timeout, could you check if the memory jumps correlate with times you're pushing new builds or scaling replicas? It might not be the route count, but the rate of change behind them.
The deployment schedule correlation is a critical diagnostic step that gets overlooked. I've plotted router memory against our CI/CD pipeline timestamps before and found nearly perfect alignment with rolling updates.
Your point about the garbage collector not getting a "clean shot" is exactly right. In our case, the memory would drop slightly after a reload, but never back to the previous baseline, creating that staircase. The new config blob allocates fresh memory before the old one is fully collected, and if another reload queues up, pressure increases.
This is where monitoring the `router_reloads_total` metric rate is more telling than the template size alone. A cluster with frequent scaling might benefit more from tweaking `ROUTER_RELOAD_INTERVAL` to introduce a short cooldown, rather than just adjusting DNS or cache parameters. It batches endpoint changes, giving GC a window to catch up.
Garbage in, garbage out.
Yeah, the connection creep is a classic silent killer. I've been burned by that `router_internal_upstream_count` metric before, though - it can be a bit misleading if you have rapid connection churn from health checks or short-lived API calls. The key is to graph it alongside the actual memory usage and look for a persistent baseline that only goes up, not the spikes.
One trick that helped us pin it down was adding a unique request ID header in our app and checking the router logs for its lifespan. We found a batch of "zombie" connections that weren't showing up as active in the main metrics but were still holding buffers open because of a keepalive misconfiguration downstream.
What's your monitoring stack look like? Are you able to correlate the router's connection states (like FIN-WAIT) with the specific backend pods?
Integration Ian