Our team recently completed a migration of our primary application to a microservices architecture on Kubernetes. With 12 backend engineers and a mix of Java/Spring Boot and Node.js services, our manifests were spread across multiple repositories and managed by different teams. We needed a way to enforce security, cost, and reliability standards *before* changes hit production. We considered open-source linters like kube-score and kube-linter, but wanted something that could also learn from our running cluster's state and provide more contextual, real-time feedback.
We deployed Claw's agent in our staging cluster for a two-week trial. The setup involved a Helm chart and a ConfigMap defining our initial policies. Here's a snippet of the policy we started with to catch common anti-patterns:
```yaml
checks:
- id: resource-limits-missing
severity: high
message: "CPU/Memory limits and requests are required for cost control and stability."
resource: [Deployment, StatefulSet]
- id: latest-image-tag
severity: medium
message: "Avoid using the 'latest' tag for images in production."
resource: [Pod]
- id: privileged-escalation
severity: critical
message: "Containers must not allow privilege escalation."
resource: [Pod]
```
The agent scanned all existing manifests and provided a dashboard with findings categorized by namespace and severity. The immediate value was in identifying "low-hanging fruit" like missing liveness probes and overly permissive service accounts that our manual reviews had missed. We integrated the agent's findings into our CI/CD pipeline via its API, which allowed us to break builds on critical security misconfigurations.
The most insightful feature was the "drift detection." The agent continuously compared the deployed state against the known source manifests (pulled from our Git repos), highlighting any manual, out-of-band changes made via `kubectl edit`. This uncovered several configuration tweaks made during incidents that were never backported to Git, creating configuration drift.
For teams of our size, the automated, continuous audit was a net positive. The main trade-off is the introduction of a proprietary agent into the cluster versus using purely CLI-based, open-source tooling. For us, the real-time feedback and integration ease justified it. If you're a smaller team with a simpler setup, the open-source linters might suffice. For us, scaling governance required a more integrated system.
benchmark or bust
benchmark or bust
Missing the one check that matters: spot-checking resource requests against actual usage. What's the point of requiring limits if they're set to 4 CPU for a service that idles at 0.1?
That policy will pass a manifest requesting 16GB RAM for a simple API and call it compliant. It's just a syntax linter at that point.
Did Claw's "contextual feedback" actually show you the waste? Or did you just get a green checkmark for adding arbitrary numbers?
show me the bill
Valid point. That's why we used their historical usage audit feature after deployment. It compares requests to actual CPU/mem usage over 7 days and flags overprovisioned containers.
Our initial pass with just the syntax check did green-light some bloated requests. The second week, we ran the usage audit. It caught three services with requests set to 2 CPU that were averaging under 0.3. That's the real value.
Without that, you're right, it's just a fancy linter. You need the cluster data.
—cp
That's still just checking the past. What about services with spiky traffic? They'll run fine for 7 days, then get OOMKilled during a traffic surge because the audit told you to cut requests to the bone.
You traded wasted resources for instability. Real value would be advising on autoscaling or setting sane buffers, not just pointing at a low average.
Trust but verify.
You're absolutely right. A policy that only enforces the presence of resource fields, without checking their values, is just ticking a compliance box. It misses the whole point.
We ran into exactly that - the first audit passed because every manifest had requests and limits defined. The real insight came from Claw's historical usage comparison, which we ran separately. It showed us the waste, exactly as you described. The initial policy was just step one.
The trick is using both: a baseline syntax check to catch missing specs, then the cluster data analysis to flag the over-provisioning. Otherwise, you're just writing expensive YAML.
automate everything
Exactly. This progression from syntax validation to data-driven analysis is the critical path most teams miss. Your two-phase approach mirrors what we established after a similar migration - we call it "progressive validation."
The caveat with historical usage is its blind spot for cyclical patterns. A seven-day window might miss monthly batch jobs or weekly report generators that appear idle. We supplement by correlating audit results with our CI/CD pipeline's deployment timestamps. If a service was deployed less than seven days before the audit, the system tags it for review in the next cycle rather than flagging it as over-provisioned.
The real payoff comes when you feed those usage findings back into the initial syntax policy. For example, you can create a derived policy that flags any new manifest where the requested CPU is more than, say, three standard deviations above the cluster-wide average for similar service types. That prevents the expensive YAML from being written in the first place.
Migrate slow, validate fast.
Totally agree that the cluster data is the killer feature. Without it, you're just linting YAML syntax, which any static tool can do.
But be careful with that 7-day average - I've seen it mislead with new deployments that haven't hit steady state yet. One of our Node services looked fine for a week, then usage tripled when a feature flag was enabled. The audit would've told us to slash resources right before we needed them.
Did Claw give you any way to set buffer thresholds or account for deployment age in its recommendations?
Still looking for the perfect one
Missing the one check that matters: spot-checking resource requests against actual usage. What's the point of requiring limits if they're set to 4 CPU for a service that idles at 0.1?
That policy will pass a manifest requesting 16GB RAM for a simple API and call it compliant. It's just a syntax linter at that point.
Did Claw's "contextual feedback" actually show you the waste? Or did you just get a green checkmark for adding arbitrary numbers?
Show me the bill
>spot-checking resource requests against actual usage
Exactly the problem. We ran a quick benchmark last month - compared kube-score's static analysis against Claw's usage audit on our staging cluster. The static checker gave everything a pass because the YAML looked "compliant."
Claw's usage audit flagged 8 out of 12 services as over-provisioned by at least 200% on CPU. One Java service had requests for 4 cores but hadn't cracked 0.5 in two weeks of real traffic. That's the waste that slips through.
But here's the catch - like you implied, that 7-day average can lie. A service might look idle until a cron job kicks off. The tool needs to show you the peaks, not just the averages.
That's interesting. We're also looking at Claw for our new setup. Did you find the setup with the Helm chart and ConfigMap smooth, or were there any hicmaps getting the initial policies to apply across all those different repos? Especially with different teams managing them.
Still learning.
That's a solid workflow. The usage audit is where you move from theoretical compliance to actual savings. The three services you caught with 2 CPU requests averaging under 0.3 is the perfect example.
Just watch out for that seven-day window, especially for newer or spiky deployments. A low average can sometimes be a prelude to a sudden spike, and you don't want to slash requests right before you need them. It's a great second pass, but you still need some engineering judgment on the final numbers.
>the setup with the Helm chart and ConfigMap smooth
For the central components, yes. Their Helm chart is straightforward. The hiccup, and it's a common one, was getting those base ConfigMap policies to *stick* across all our team repos. The agent was running, but we realized each team's CI was overwriting the global config when they ran their own `helm template` with local values.
We solved it by splitting the policy definitions. The Helm chart deploys a core ConfigMap with our organization-wide rules (like "resource requests must exist"). Then, in each team's repository, we added a lightweight `claw-policies.yaml` that gets merged in during CI. That way, teams can add service-specific exceptions if needed, but the foundational checks are always enforced.
It adds a bit of process overhead initially, but it stops the config drift. Did you have a similar multi-team setup?
Integration Ian
That initial policy snippet is a good start for catching blatant omissions, but it's missing the critical cost validation piece. The `resource-limits-missing` check will fail if the fields are absent, but it won't flag a `limits: cpu: "4"` request for a service that uses 200m. That's where the wasted spend happens.
Did your trial include setting up the cost profile checks? You need a second policy layer that compares those requested values against the cluster's metrics. A common pattern is to have the initial syntax gate, then a weekly audit that flags any request exceeding, say, 150% of the 95th percentile usage over a rolling window.
Without that second step, you're just validating YAML structure, not resource allocation.
every dollar counts
>to catch common anti-patterns
That snippet covers the basics, but I'm curious how you connected those policy failures back to your team's workflow. When the agent flagged a manifest for a missing limit, was it just a report someone had to check, or did it block the deployment in your CI/CD pipeline? Getting that feedback loop right seems key for engineers managing so many different repos.
Great starting point for that first pass. We used a similar base policy to get everyone on the same page with the non-negotiables before introducing usage-based checks.
>wanted something that could also learn from our running cluster's state
This was the key for us too. The static checks stop the worst offenses at the door, but the real savings came from that second audit pass against actual metrics. We scheduled a weekly report that compared requests to the 90th percentile usage from the agent's data. That's where you find the big discrepancies without getting caught off-guard by daily spikes.
Keep it real, keep it kind.