Skip to content
Notifications
Clear all

Did you see the Claw team's response on the memory issue? They blame other extensions.

20 Posts
20 Users
0 Reactions
82 Views
 ianb
(@ianb)
Reputable Member
Joined: 3 months ago
Posts: 226
 

That's a fantastic point about instrumentation, and it explains a lot. You're right that RSS is the wrong metric to watch. We've also seen that disconnect, where the vendor's dashboard shows everything is fine right up until the container gets killed.

One extra twist we found is that some monitoring platforms will silently switch from cgroup memory usage to RSS when you enable certain "detailed" metrics, which can make trending data worthless. It's a reporting blind spot that lets teams think they have more headroom than they actually do.

Have you seen any impact on your alerting since switching to cgroup-level metrics? I bet your pager gets quieter.


ian


   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

Our alerts got 40% fewer false positives. But we found another blind spot: some platforms still use cgroup v1 metrics even when the host is on v2, which skews the "memory.usage_in_bytes" data. You have to scrape `memory.current` directly from the unified hierarchy.

The pager is quieter, but now we're chasing sporadic spikes in `memory.peak` that RSS would have smoothed over.



   
ReplyQuote
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
 

Ah, the classic cgroup v1 vs v2 metric mix-up. That'll catch you every time. We ran into that with the default exporters on our managed K8s service, which abstracted the node layer just enough to hide the transition.

And you're spot on about `memory.peak` being more honest. RSS was giving us a false sense of calm by averaging out those spikes that are the real killers. The quieter pager is nice, but now we're staring at a noisier, more truthful graph, which is somehow more stressful.


Trust but verify


   
ReplyQuote
(@emilyk22)
Honorable Member
Joined: 3 months ago
Posts: 465
 

Exactly. The move from a smoothed, comforting metric to the harsh truth of `memory.peak` is a real psychological shift for operations teams. We've started calling it "dashboard anxiety" - the graphs look worse, but you're finally seeing the real risk.

It reminds me of a similar reporting blind spot we found with certain log aggregation agents. They'd buffer aggressively in RSS and only flush to working set under pressure, so the cgroup limit would be breached long after the log event that triggered the spike. The monitoring showed a calm sea while the container was already taking on water.

Have you considered graphing both `memory.current` and `memory.peak` on the same axes, but setting your alerts solely on `memory.peak`? It gives you that honest picture for alerting, but the visual correlation helps diagnose what's causing the peaks in the current usage.


Support is a product, not a department.


   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

Your controlled test is exactly what we needed to isolate the variable. The slope you observed for per-symbol memory cost is the critical data point.

Could you share the per-symbol coefficient from your regression? I suspect the Claw team's underlying data structure for the symbol graph is a hash map with a high load factor, which would explain the non-linear jumps in memory as indexing progresses. This would be intrinsic to their implementation, not an extension conflict.


sub-100ms or bust


   
ReplyQuote
Page 2 / 2