Skip to content
Notifications
Clear all

Thoughts on the new Boundary worker performance metrics?

25 Posts
23 Users
0 Reactions
71 Views
(@davidh)
Honorable Member
Joined: 3 months ago
Posts: 410
 

I agree that the excitement is justified. Moving from a black box to any instrumentation creates a tangible operational mindset shift. The lack of histograms forces us into the correlative detective work you're describing, but that process itself often uncovers other systemic dependencies you'd otherwise miss.

On the scrape noise, setting a longer interval is a good start, but you also need to be selective with what you actually ingest. The Go runtime metrics are the usual suspects. In our environment, we had to aggressively drop high-cardinality `go_memstats_alloc_bytes` series per pod after the first scrape, as they were bloating our TSDB. The `process_` metrics are also redundant if you're already scraping the node exporter. It's a balancing act to capture the new Boundary signals without the Go telemetry drowning them out.

One subtle quirk we found: the initial scrape after a worker restart can show artificially high rates for the first interval due to counter initialization, which can trip naive alerts. Using `rate(metric[5m])` with an offset of at least two scrape intervals in your queries helps avoid that false signal.


Data over dogma


   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

Totally agree on the scrape interval. We started at 15s and those Go metrics were definitely contributing to noisy graphs. Moving to 30s helped a lot with that baseline chatter.

The correlative approach with active connections and audit logs is clever, but I found the timing is often just a bit too loose for catching brief latency spikes. It's great for long-term trend analysis, but less so for pinpointing a real-time issue.

Durable storage for the counters would be nice for absolute totals, but honestly, I've found that focusing on the rates and that bandwidth limit percent gives me a much clearer picture of current system health. It forces you away from those "lifetime total" vanity metrics.


✌️


   
ReplyQuote
(@henryf)
Reputable Member
Joined: 3 months ago
Posts: 291
 

You're right about the scrape interval, but the real noise culprit is the default Go metrics. I'd drop them entirely. If you need the data, pull it from the node exporter instead.

Focusing on rates over totals is the only sane way to run Prometheus. "Lifetime totals" are meaningless for anything but a static report. The bandwidth limit percent is the real metric, everything else is just a derivative signal.



   
ReplyQuote
(@annas)
Honorable Member
Joined: 3 months ago
Posts: 542
 

Dropping the Go metrics entirely is the pragmatic move, I've done the same in production. The node exporter gives you a cleaner, normalized view of the process anyway. You just have to be careful to keep the `go_threads` metric from Boundary's endpoint if you're not already getting it elsewhere, as it's a decent canary for goroutine leaks in the worker's proxy routines.

On rates being the only sane way, I mostly agree, but that blanket statement has a real cost during post-mortems. When you're trying to reconstruct total data transfer for a specific compliance window or a cost attribution report, you're stuck with log aggregation because the counter resets are lost. It's a trade-off the design makes, but calling lifetime totals meaningless ignores the paperwork part of the job.

The bandwidth limit percent is indeed the primary signal. But in my deployments, it's been a lagging indicator, not a leading one. By the time it climbs, the upstream churn in `active_connections` has already been spiking for minutes. You need to watch the derivative of the rate of connection changes to get ahead of it.



   
ReplyQuote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

Good point on the paperwork angle. Compliance windows don't care about your elegant rate calculations.

But that lagging indicator problem you mentioned is the real issue. Watching the derivative of connection rates is fine, but it's reactive. If `active_connections` is already spiking, the session setup process is already under stress. The leading indicator I look for is a sustained drop in `go_threads` while `active_connections` climbs - suggests the proxy routines are failing to spawn, a classic pre-churn signal.

You're keeping `go_threads` from the endpoint. How's the cardinality on that been for you? I've seen it create a mess when workers are cycled frequently.



   
ReplyQuote
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
 

That go_threads vs active_connections divergence is a solid catch. It's a better leading signal than connection rate derivatives.

The cardinality on go_threads has been fine for us, but our workers are long-lived pets. If you're cycling workers frequently in an autoscaling group, each new pod creates a fresh series. That's where you'd want to use a recording rule to strip the pod label and aggregate, or drop it entirely and rely on a separate runtime-level dashboard.


Trust but verify, then don't trust.


   
ReplyQuote
(@contrarian_coder)
Reputable Member
Joined: 7 months ago
Posts: 309
 

That bandwidth limit percent alert is the right idea, but 80% feels arbitrary without knowing your burst profile. I've seen workers hover at 90% for hours during sustained transfer jobs without issue, then spike to failure from a cold start at 60% when session churn kicks in.

Correlating active_connections with go_goroutines for stalled teardowns is clever, but it's already a failure state. By the time you see that lag, sessions are stuck. You'd be better off watching the rate of change on connections versus a baseline. A linear climb in connections with a flat goroutine count is the quieter signal that the pool is starving.


prove it to me


   
ReplyQuote
(@emma78)
Reputable Member
Joined: 3 months ago
Posts: 221
 

Good to see the metrics endpoint is getting more useful. I'm still getting started with Boundary monitoring.

You mentioned building a dashboard for session establishment latency. Are you calculating that just from the active connections metric, or are you pulling in other data? I'm trying to figure out what a realistic baseline for my own setup should be.

Also, for someone focusing on the health of the overall system rather than deep debugging, which one or two of the new boundary_ prefixed metrics would you watch first?



   
ReplyQuote
(@alexw)
Reputable Member
Joined: 3 months ago
Posts: 443
 

For a baseline on session establishment latency, you can't get it from the active connections metric alone, that only tells you how many are alive at a given moment. You need the `boundary_session_client_active_session_count` metric and its derivatives to see the rate of session creations over time. Correlate that with timing data from your audit logs to build a latency picture.

If you're focusing on overall health, start with `boundary_proxy_bandwidth_limit_percent` and `boundary_session_client_active_session_count`. Watch the percent for sustained saturation near your configured limit and the session count for unexpected drops, which often happen before latency becomes apparent.

Your baseline will be unique to your network and target types, so start by graphing those for a week during normal operations to see your own pattern.


Stay grounded, stay skeptical.


   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

The `bytes_upstream` counters reset on worker restart. They're stored in memory only. This is a known limitation for any pure counter in the Go Prometheus client.

I've seen session establishment latency best tracked indirectly. You're right about the lack of histograms. The workaround is to calculate the delta in `boundary_session_client_active_session_count` over a short window and correlate with timestamps in the audit logs. It's not elegant, but it gets you a baseline.

One gotcha: the `boundary_worker_proxy_active_connections` metric labels include the target protocol. If you're using a mix of SSH, RDP, and TCP, the cardinality can jump. It's manageable, but you'll want to aggregate by protocol in your recording rules from the start.



   
ReplyQuote
Page 2 / 2