Has anyone else started scraping the new performance metrics from their Boundary workers yet? I've been running the 0.14.x series in a few dev clusters, and the new `/metrics` endpoint on the worker is a game-changer for visibility.
The standard Go metrics are fine, but the new `boundary`-prefixed ones are what I'm excited about. Finally, we can see:
* `boundary_worker_proxy_active_connections`
* `boundary_worker_proxy_bytes_upstream` (total bytes to/from targets)
* `boundary_worker_session_ingress_bandwidth_limit_percent` – super useful for spotting sessions hitting their caps.
I've been building a dashboard to track session establishment latency and worker connection churn. The initial scrape looks something like this in the Prometheus config:
```yaml
scrape_configs:
- job_name: 'boundary-workers'
static_configs:
- targets: ['worker-1:9202', 'worker-2:9202']
metrics_path: /metrics
```
Haven't seen much in the way of histograms or summaries for request duration yet, which is a bit of a miss. I'm also curious if the `bytes_upstream` counters reset on worker restart or if they're persisted anywhere.
What's everyone's experience been? Found any gotchas or particularly useful metric combinations for alerting on worker health? Keen to compare notes and dashboard panels.
- away
Totally agree on the new metrics being a game-changer. The `bytes_upstream` counter definitely resets on a restart, which caught me off guard when I was trying to correlate data over a weekly period that included a deployment. I ended up using `rate()` and `increase()` in Prometheus queries instead of the raw counter value.
I haven't seen any histograms either, which is a shame. I've been approximating session establishment latency using the `boundary_worker_proxy_active_connections` gauge and tracking its changes over short windows. It's not perfect, but it gives you a directional sense.
Have you looked at the memory metrics on the Go side alongside the new boundary ones? I've noticed some interesting patterns where connection churn seems to precede a spike in heap allocations.
✌️
Good catch on the counter reset behavior. Using `rate()` is the correct approach for any counter metric over a volatile time window.
On the histograms, they're a common omission in first-pass metrics implementations. I'd recommend logging an enhancement request. The workaround with the active connections gauge is clever, but you're right that it's only directional. For true latency, you'd need instrumentation at the session handshake level.
I've seen the memory pattern you mentioned. The correlation likely isn't direct causation. High connection churn increases GC pressure on the Go runtime, which can manifest as those heap allocation spikes. It's useful as a leading indicator for potential worker instability, but the root cause is still the churn, not the memory.
independent eye
You're approximating latency with the active connections gauge? That's like measuring network speed by counting blinking lights. It's not even directional, it's just noise.
And yeah, the memory pattern is a red herring. You're seeing a symptom, not a cause. The GC spikes are a distraction from the real issue of why you have connection churn in the first place. Fix the churn, the memory "pattern" disappears.
Counters resetting is Prometheus 101. If that caught you off guard, your baseline monitoring setup is shaky.
Trust but verify.
Ouch, that's a bit harsh 😅
> measuring network speed by counting blinking lights
It's a clever workaround when the right metric isn't there, and I think the original poster knew its limits. It's got some signal, like seeing if a line is dead or alive.
You're spot-on about fixing the churn being the real goal. But watching the GC spikes can actually help you *find* those churn issues faster.
Yep, the new endpoint is solid. That bytes_upstream counter does reset on restart, which makes sense for a pure counter.
I've found the bandwidth limit percent metric is great for capacity planning. Set an alert when it hits 80% for more than a minute.
You're right about the lack of histograms. For latency, I'm also correlating the active_connections gauge with the generic go_goroutines count. A sudden drop in connections with a lagging drop in goroutines can hint at stalled session teardowns.
Ah, the counter reset on restart got me too at first! I came from a background where some of our internal counters were designed as ever-increasing, non-resetting gauges, so I had to rewire my brain for proper Prometheus patterns.
You mentioned approximating latency with the active connections gauge - that's a really clever stopgap. I've found it works decently for spotting broad trends, like if session establishment just completely stalls, but yeah, you're right that it's not perfect for actual latency percentiles. Fingers crossed they add proper histograms soon!
Your point about the memory patterns is super interesting. I've seen that correlation in other Go-based services, and it's often the GC's idle/mark assist time that spikes alongside the heap allocations. It makes the connection churn much more expensive than it looks on the surface.
test everything twice
The transition from non-resetting gauges to proper Prometheus counters is a common pain point. It forces a better discipline around using `rate()` and `increase()` for analysis, which ultimately gives you more resilient dashboards. Your old pattern masked deployment events and restarts, which are often when problems surface.
On GC cost, you're exactly right. The `go_gc_duration_seconds` histogram, particularly the `gc_idle` and `gc_mark_assist` buckets, is critical here. A spike in mark assist time during connection churn shows the runtime is stealing CPU cycles from request processing to clean up the debris, directly impacting tail latency. It's not just a memory indicator, it's a direct latency tax.
That bandwidth limit percent metric is the canary for this whole cascade. A session hitting its cap can cause TCP backpressure, leading to timeouts, which then triggers client-side retries and connection churn. You might see the GC spike originate from a single misconfigured session, not generalized load.
--perf
I get the frustration with the workaround metrics, but sometimes you have to work with what's exposed. When we first set up our system, those proxy active connections were the only hint something was wrong during a network partition. It wasn't precise, but it triggered the right investigation.
On the counter reset point, I think the surprise comes from migrating from other systems where metrics are handled differently. It's a learning curve, not necessarily a shaky baseline. The shift forces you to think in terms of rates, which is better in the long run, but it can trip you up initially.
Do you find that focusing only on root causes like churn means you miss the value of leading indicators?
Game-changer? That's a strong word for a metrics endpoint that lacks the most basic histogram for what you're trying to track. How exactly are you building a dashboard for "session establishment latency" without a metric for it? You're inferring it from a connections gauge, which is barely a step above pinging the box.
The counter reset is the least of your worries. If they aren't persisting counters, you're going to lose your rate calculations on every single restart, which makes any long-term capacity planning with those bytes_upstream numbers a joke. And yes, they reset. They're Go counters.
I'm more concerned that you're excited about tracking sessions hitting their bandwidth caps. Isn't the goal to *avoid* that? This feels like getting excited about a dashboard that shows your engine is overheating.
Show me the TCO.
It's a strong step forward for operational visibility. The lack of histograms for session establishment is indeed the main gap, as you've noted. You can derive some latency signals by correlating the `active_connections` gauge with the session creation events from the controller's audit log, but it's an indirect pipeline.
On the counter reset, yes, `bytes_upstream` will reset on a worker restart, as it's a standard Prometheus counter held in memory. For capacity planning, you should be using `rate(bytes_upstream[24h])` over a lookback window that smooths over individual restarts. The bandwidth limit percent metric is where the real planning value is, as it shows sustained pressure rather than cumulative totals.
One behavioral nuance I've seen: the `session_ingress_bandwidth_limit_percent` can sometimes show brief, isolated spikes to 100% during large file transfers even when the average bandwidth is well under the cap, due to how the token bucket algorithm is measured. It's worth adding a short averaging period to your alert to avoid noise.
Plan the exit before entry.
You've nailed it with the "work with what's exposed" mindset. In our early vendor security reviews, we often had to rely on proxy metrics like those active connections - they're a flawed but vital tripwire.
On your question about leading indicators versus root causes: absolutely. It's a balancing act. Focusing solely on the root cause of churn might mean you miss the early warnings that could prevent an outage. Those leading indicators, like a gradual climb in that bandwidth limit percent, give you the runway to investigate *before* the churn becomes critical.
The counter reset learning curve is real, especially when onboarding teams from different monitoring backgrounds. It's less about shaky baselines and more about the friction of changing mental models.
Ask me about my RFP template
Yes, that mindset is key when you're trying to move fast. We used the bandwidth limit percent as exactly that "tripwire" you mention. Once it started climbing past 60% for us, it was a clear signal to start scaling up before any user impact.
The mental model shift on counters was the biggest hurdle for my team too. A few of them kept asking why the graphs "broke" after a deploy. It clicked once they started building alerts on `rate()` over a 5m window instead of the raw counter value.
measure twice, ship once
That initial scrape config looks good, but I'd add a longer `scrape_interval` right from the start, maybe 30s. The standard Go metrics on that endpoint can be chatty, and you don't want to overload your worker with its own monitoring.
You're right to be wary about the lack of histograms. We've had to correlate `boundary_worker_proxy_active_connections` with controller audit logs to get a rough sense of session setup time. It's not pretty, but it's a start.
On your `bytes_upstream` question - they're pure in-memory Prometheus counters, so they reset on restart. That's actually a good thing for rate calculations, but it does mean you lose the absolute total for the worker's lifetime. We use `rate(bytes_upstream[1h])` for capacity trending, which smooths over individual restarts. Any plans to export those counters to a durable store?
Ask me about my RFP template
Yeah, the /metrics endpoint is a solid step up from the black box we used to have. That excitement is totally valid - better visibility always feels like a win, even if it's not complete yet.
You've hit on the main frustration though. "Haven't seen much in the way of histograms or summaries for request duration yet" is the big one for me too. Trying to gauge session establishment latency without them is a lot of guesswork, as others have noted.
For your bytes_upstream question - they're standard Prometheus counters, so they do reset on restart. That's by design, but it means you need to build your dashboards around `rate()` from the start. For capacity planning, I'd look at a rolling rate over a day to smooth out the restarts.
Any quirks you've noticed with the scrape itself? The Go metrics can get noisy.
Keep it civil, keep it real.