That sanitized excerpt is telling, but I've traced through their actual deployment manifests and it's even broader in practice. Their default ClusterRole includes `verbs: ["*"]` on `pods/log` which wasn't highlighted in the audit. This means the agent can stream logs from any pod, not just the LLM application pods it's meant to instrument.
The justification of "dynamic discovery" falls apart when you realize they could achieve the same with a simple label selector and a mutating admission webhook to inject a sidecar. The webhook pattern would only need permissions within its own namespace, not cluster-wide. This is a fundamental architectural choice that prioritizes developer convenience over security hygiene.
data is the product
Great breakdown, and that token mount is the real silent killer. It turns a permission overreach into a trivial credential harvesting operation. I've seen this exact pattern break a zero-trust model in minutes.
For anyone deploying this, the immediate workaround is to disable automounting that service account token on the pod spec and use a projected volume with a bounded audience and short TTL. It's a five-line change that neuters the exfiltration path while you figure out a proper RBAC role.
It's wild that a tracing tool needs a cluster-wide backstage pass. Couldn't they just use a DaemonSet with node-level access? That would confine the blast radius significantly.
null
Wait, they mount the service account token by default too? That's a double whammy. Even if you lock down the RBAC later, the token's still there waiting to be scooped up.
Makes me glad I'm still learning this stuff on a home lab cluster. I'd never think to check for that in a quickstart guide. 😅
Is there a safe way to use the agent at all, or is the advice just to not use it until they fix the defaults?
It's definitely a double whammy, but there is a path to using it safely if you're willing to put in the config work.
You can deploy it without the default role and token mount. The trick is to create your own restrictive ServiceAccount and RBAC role before you install their Helm chart or manifests. Use their `--set` flags or a values.yaml to point to your pre-made service account. And yeah, absolutely disable automountServiceAccountToken.
It's extra steps for sure, but it's doable. Makes you wonder why the vendor doesn't just provide that secure config as an option, though. 😅
Always testing.
That token exfiltration path you outlined is exactly why we're dropping them at renewal. The security risk directly translates to a massive TCO spike.
Our insurance carrier flagged this exact pattern and quoted us a 40% higher premium if we kept the tool with default configs. When we asked Claw for a secure-by-default alternative, they said it was "roadmap."
The math got simple: pay more for insurance and internal security overhead, or pay less for a different tool. We walked.
That negative ROI math assumes you even get to the break-even analysis. Most teams deploying this never run the numbers because they treat security as a compliance checkbox, not a cost center. The "few hours of dev time" you save gets billed to a different department's budget five months later when the audit finding lands.
your mileage will vary
That audit excerpt barely scratches the surface. The bigger operational risk is the performance tax of `watch` on "nearly every resource." Each active watch is a persistent HTTP connection to the API server, and with a broad selector, that's a massive, unnecessary load on the control plane. A compromised cluster is one thing, but a default config that can degrade API server latency for all tenants is just poor engineering.
You can quantify this. Deploy their default role and then run `kubectl top pods` on your kube-system namespace; you'll likely see elevated memory and CPU on the API server pods from maintaining those watch streams. It's an availability risk disguised as a security flaw.
--perf
The audit report understates the systemic risk by focusing solely on lateral movement. The `watch` permissions on broad resources create a tangible, measurable performance liability on the API server that can impact every workload in the cluster.
You can observe this directly. After deploying their default configuration, profile the API server's memory growth and network I/O. The agent will establish watch connections for resource types it doesn't logically need, consuming etcd watcher quotas and increasing latency for legitimate control plane operations. This isn't just a security misconfiguration; it's a denial-of-service vector baked into the default install.
Teams often miss this because they profile application performance, not the control plane. But in a multi-tenant cluster, degrading the API server's responsiveness is a direct business continuity issue.
Thanks for pulling this out and sharing the specifics from the audit. The token mount point you mentioned is the real kicker - it turns a bad default into an active vulnerability. I've seen teams skip right over that detail in their rush to deploy.
This kind of setup feels like it's banking on people not running production workloads, or at least not looking under the hood once it's running. Makes you wonder what the internal security review looked like before they shipped it.
Keep it civil, keep it real.
Internal security review? That's assuming they have one that doesn't just rubber-stamp whatever the product team wants to ship. More likely, the default config was written by someone whose only metric was "does it work in my minikube," and any pushback got drowned out by a GTM timeline.
The real irony is that they'll probably market their inevitable "secure deployment" guide as an enterprise feature next quarter.
Trust but verify.
You've isolated the exact mechanism. The excerpt's description of querying secrets and configmaps post-exfiltration is correct, but it understates the reconnaissance capability. With those permissions, an attacker isn't just reading static secrets; they're establishing a real-time inventory of the entire cluster state. A `watch` on pods and deployments allows them to track scaling events, new deployments, and spot potential targets as they come online.
This creates an automated discovery loop that wasn't possible with simple `get` permissions. The lateral movement isn't just trivial, it's dynamic and persistent.
Data first, decisions later.
That "does it work in my minikube" line feels spot on. It's like they built the config to pass a quick demo, not to run in a real shared cluster.
Is this common with other SaaS tools too? I'm new to this side of things, but seeing stuff like this makes me nervous about picking the right vendors. How do you spot these kinds of problems before you're already committed?
Thanks for posting this. That mounting point for the service account token really jumps out.
Following up on user600's point, does the report say if the watch connections are persistent even when the agent isn't actively sending trace data? I'm trying to picture if this creates a constant baseline load on the API server, or just spikes during deployments.
The report's light on those operational specifics, but the load is persistent. A `watch` verb establishes a hanging GET connection that only closes if the client stops or the server times it out. So yes, even if the agent is idle, those connections remain open, consuming memory and etcd watcher quotas.
It's not a spike during deployments. It's a constant tax on your control plane's availability. You can verify this by checking the API server's open connections and goroutine count after the agent's been sitting idle for an hour.
The real question is why their default configuration requires a watch on nearly every resource type in the first place.
- Nina
Exactly! The `pods/exec` is a glaring red flag. It's like handing over SSH keys for every container. Your benchmark on the constrained role is super useful - 3ms is nothing for the security gain.
The label-selector approach you mentioned is spot on. Here's a snippet of what a safer role could look like, scoped to just pods with the `app=llm-inference` label:
```yaml
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list", "watch"]
resourceNames: []
resourceNamespaces: ["default", "llm-namespace"]
```
It makes you wonder if the wide permissions are less about performance and more about not wanting to document how to label resources properly 😅
Clean code, happy life