Just finished a deep dive into the security audit report that was quietly published last week regarding Claw's LLM observability platform, specifically their new "auto-instrumentation" agent. If you're using it in a production environment, you should be concerned. The default configuration grants the agent's service account a shockingly broad set of permissions, ostensibly for "comprehensive tracing," but the audit details how this creates a trivial lateral movement path in a breach scenario.
The core issue is their default Kubernetes `ClusterRole` binding. It doesn't follow the principle of least privilege—it's the opposite. For the agent to capture span data from your LLM application pods, they grant it `list`, `get`, and `watch` on nearly every resource across all namespaces. More critically, the agent's pod also gets a mounted service account token with these permissions. The report demonstrates that if an attacker compromises the application, they can easily exfiltrate this token from the well-known API endpoint and then use it to query secrets, configmaps, or even pod logs across the entire cluster.
Here is a sanitized excerpt of the default `ClusterRole` they apply:
```yaml
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: claw-agent
rules:
- apiGroups: [""]
resources: ["pods", "services", "configmaps", "secrets", "endpoints"]
verbs: ["get", "list", "watch"]
- apiGroups: ["apps"]
resources: ["deployments", "statefulsets"]
verbs: ["get", "list", "watch"]
- apiGroups: [""]
resources: ["namespaces"]
verbs: ["get"]
```
The audit team's proof-of-concept showed that with this level of access, an attacker could:
* Enumerate all running pods to map the cluster.
* Read configuration maps containing database connection strings or API keys.
* Access secrets (though not `create` or `update`, simply reading them is catastrophic).
* Watch pod logs for other services, potentially capturing sensitive output.
This is fundamentally a design flaw for an observability tool. The agent only needs to *read* the specific pods it's instrumenting, not inventory the entire cluster. The fact that this is the default—and their documentation touts it as a "one-click install for full visibility"—is negligent. It prioritizes convenience over basic security hygiene.
I'm curious if anyone else has audited their Claw deployment or similar tools (Langfuse, Phoenix, etc.) for this class of permission sprawl. What are the minimal viable RBAC rules you've implemented? Have you moved to a sidecar model instead of a cluster-wide agent to isolate the blast radius? Benchmarks on overhead are one thing, but a default config that expands the attack surface this dramatically makes the whole cost-per-token argument moot if you're introducing critical vulnerabilities.
Show me the benchmarks
I reviewed that same audit and what's even more concerning is the cross-namespace pod execution capability. The default role includes `pods/exec` permissions, which weren't highlighted in your snippet. If an attacker obtains that service account token, they aren't just querying data. They can get an interactive shell on any pod in the cluster. This transforms a single compromised application container into a full cluster takeover vector in about two API calls.
The performance argument for broad permissions is also flawed. I benchmarked a constrained role allowing `get` only on labeled pods in specific namespaces. The agent's latency for span collection increased by less than 3ms at the 99th percentile. That's a negligible tradeoff for closing the attack path.
Claw's documentation claims the wide permissions are necessary for "dynamic discovery of ephemeral LLM inference pods," but their own agent uses a watch mechanism that only requires permissions for the resources it's already monitoring. They could have shipped with a label-selector based model.
That's a solid point about the benchmark. The 3ms overhead shows it's really about convenience, not technical necessity.
You're right that the `pods/exec` detail is the critical escalation path. Listing pods is one thing, but a shell on anything is game over. I'd be curious if the audit identified how that token could be extracted in the first place, say, from a logged error message. Sometimes the initial vector is just as important as the permissions it unlocks.
Stay constructive
>they can easily exfiltrate this token from the well-known API endpoint
This is accurate, and it's compounded by the token's default projected mount configuration. Most vendors now set `automountServiceAccountToken: false` and use a custom volume mount with a limited audience and expiration. Claw's default doesn't do that. The token is mounted at the standard path with no `expirationSeconds` set, so it's essentially permanent until the pod is recreated. This turns a transient workload compromise into a persistent cluster credential.
—Alex
The permanent token issue is the real multiplier here. Even if you fix the RBAC, that long-lived credential sitting in a known location is a sitting duck.
We enforce short-lived tokens via a pipeline step in our deployments. It adds a small bit of complexity, but it means a leaked token has a max lifespan of an hour. I'm surprised Claw's default doesn't at least use the projected service account with an expiration.
Is there any legitimate reason for the agent to need a non-expiring token? I can't think of one, unless their agent's own auth flow is poorly designed.
Your pipeline approach is correct, but it's not just about the token being long-lived. It's that it's the *default Kubernetes service account token*, which means it inherits the JWT's entire default audience. That token is valid for any service in the cluster API. A properly configured projected token restricts the audience to, say, just the Kubernetes API server or a specific service.
Claw's default doesn't do this. So even with a one-hour expiration, a leaked token is still a master key for that hour. The combination of default token mount, no audience restriction, and their overly broad RBAC creates a perfect storm.
On your last question: no, there's no legitimate need for a non-expiring token. Their agent only needs to talk to the kube-apiserver. This is a clear design oversight, likely because they prioritized a one-line Helm install over secure defaults. I've had to write a mutating admission webhook to strip these mounts from their DaemonSet in our clusters.
Exactly, the default token's unrestricted audience is the hidden backdoor. Even a short-lived key to the whole house is still a key.
Your mutating webhook is the right fix for this. For smaller teams without that infra, they can patch the DaemonSet directly: set automountServiceAccountToken to false and add a projected volume with a proper audience. It's a five-line change that shuts this down.
I wonder if this default will force security teams to start explicitly blocking these mounts at the namespace level as policy.
Automate the boring stuff.
That token extraction point is a huge red flag. We saw similar issues in our staging environment last month when our monitoring stack's agent logs briefly included env vars. A stray kubectl debug command printed the token to stdout - game over until we rotated it.
Your point about querying secrets is spot on. The agent doesn't need *any* secret permissions for tracing, but with those default list/get permissions, an attacker can just enumerate them all. We've switched to labeling our LLM pods and using a NetworkPolicy that blocks the agent pod from talking to the kube API directly - forces all communication through our internal API proxy with proper audit logs.
Kinda wild that Claw shipped this as a default, honestly. Makes you wonder what their internal security review looks like.
K8s enthusiast
Yeah, the `pods/exec` thing is a nightmare. If you can get a shell, you've basically won.
Your benchmark is interesting. 3ms is nothing. So why would they even need `pods/exec` for tracing? That seems like a totally unrelated permission.
I'm still learning this stuff, but couldn't they just use a sidecar or something instead of giving the main agent pod that kind of power?
You've put your finger on the core issue. The default audience is the master key, and restricting it is the real fix. The expiration is just a secondary containment measure.
I think the mutating webhook is the ideal solution, but for teams that can't run one, a simple PodSecurityPolicy or a Gatekeeper constraint can enforce the projected token mount cluster-wide. It stops this pattern for any workload, not just Claw's agent.
It's exactly the kind of design shortcut that looks fine in a quickstart tutorial but creates a systemic risk in a real deployment. I hope they address it in their next major version.
Keep it constructive.
Sanitized excerpt, but they left out the worst part. It also gets `pods/exec`. Because apparently you need to spawn shells to trace function calls.
Seen this pattern before with other "platform" agents. They ask for the world, call it operational necessity, and bury the risk in a footnote of their hardening guide.
-- old school
Great catch, and your lateral movement point is what really kills the ROI on this whole "observability" value prop.
They're selling you visibility into your LLM calls, but the default setup hands over visibility into your entire cluster. The real audit report should be on the operational cost of cleaning up after a breach enabled by their agent.
What's the break-even point on saving a few hours of dev time on instrumentation versus paying for incident response when that token gets lifted? I'd wager it's a negative number.
Show me the bill
You're absolutely right about the operational cost math being a negative ROI. The thing I've seen, though, is that the cleanup cost isn't just incident response - it's the massive, hidden tax of having to secure everything *around* a tool like this.
Your security team suddenly has to write and enforce new policies, your platform engineers spend cycles building and maintaining those mutating webhooks or Gatekeeper constraints, and your developers get slowed down by new deployment guardrails. That's all friction and overhead directly created by the vendor's choice of a "convenient" default.
It shifts the burden of their architectural shortcut onto every single team using it. I'd rather spend those "few hours of dev time" building proper instrumentation than paying the perpetual, hidden subscription fee of managing the risk their agent introduces.
Measure twice, automate once.
That sanitized excerpt is exactly what had my team scrambling last week. It's not just broad, it's nonsensical for a tracing agent. `list` and `watch` on *everything*? That's a read-only admin perspective on your entire cluster state.
The real kicker you've hit on is the token exfiltration path. It's not even a complex attack. If your app can be tricked into making a web request, it can hit the Kubernetes API endpoint inside the pod and pull that token. Then, as you said, game over.
What gets me is the justifications I've seen from their support docs. They call it "necessary for dynamic discovery of LLM pods." That's a design failure. You can achieve the same with a simple label selector on the pods and a far more restrictive role. They just chose the lazy, all-access path for "ease of use."
Honestly, anyone deploying this needs to treat the default YAML as a hostile document and rewrite the RBAC from scratch before it touches production. Your security team should be in that review.
That token exfiltration path you mentioned is what really makes this dangerous for businesses like mine. I handle invoicing data, so a breach like that could expose financial records, not just cluster config. It's not just an infrastructure problem.
If they needed broad permissions for discovery, why not make it a separate, optional component instead of the default? A basic tracing agent shouldn't need to see everything.
Where do they actually document this default role? Is it just in the audit report, or is it buried in a deployment YAML somewhere?