Everyone's raving about the latest Claw agent, but let's be honest: half of those "performance metrics" they show you are vanity stats. I needed to see what happens when it actually chokes on a complex dependency graph or a legacy codebase with more skeletons than a graveyard. So I built a Grafana dashboard that goes beyond token counts and happy-path completions.
The core of it is tracking *latency under load* and *context degradation*. I'm not just logging average response time; I'm capturing P95 and P99 during active refactoring sessions, and correlating that with the "usefulness" score my team manually logs (simple 1-5 on whether the suggestion was directly usable). The dashboard also watches for the agent hitting context window limitsβwhen it starts silently dropping earlier parts of a large file, that's when the real bugs creep in.
Here's the key panel config for the latency heatmap. This shows you exactly when performance falls off a cliff, which usually coincides with opening a monolithic file.
```yaml
# grafana/panels/latency_heatmap.yml
panel:
title: "Claw Agent Latency vs. File Size"
targets:
- expr: |
histogram_quantile(0.95, sum(rate(claw_agent_request_duration_seconds_bucket{job="claw-backend"}[5m])) by (le, file_size_bucket))
legendFormat: "P95 Latency - File size {{file_size_bucket}}"
heatmap:
dataFormat: "tsbuckets"
yAxis:
decimals: 3
unit: "s"
```
The real eye-opener was setting up alerts for "coherence decay." If the agent's suggestions for the same function start diverging wildly in a short session (measured by a simple code AST diff score), it pings us. It's caught several instances where the agent got confused by a mid-session switch in coding style, leading to contradictory architectural advice.
Results? We've dialed back our context window from the "max everything" default to a more conservative 12k tokens for most tasks, because the dashboard clearly showed accuracy plummeting after that point, despite latency still being "acceptable." It also convinced the team to stop using the agent for large-scale, cross-file renamingβthe metrics showed a 40% increase in logical errors compared to smaller, file-local refactors. Fancy graphs won't fix your code, but they'll at least show you which tools are lying to you.
prove it to me