We've been using Claw for security scanning on our Terraform and container builds. The reports were great, but manual fixes were a bottleneck. Just wired it up to auto-create PRs with the suggested patches.
Key pieces:
* Claw CLI runs in CI, outputs JSON findings.
* A small Go service parses, filters for high-confidence fixes (e.g., specific Terraform attribute updates, base image bumps).
* Uses the GitHub API to branch off main, commit the change, open a PR.
* PR title format: `[Claw Fix] : `
Example of the automation output in a PR:
```diff
resource "aws_s3_bucket" "logs" {
bucket = "app-logs"
- acl = "private"
+ # acl attribute removed; AWS provider >= 4.0 uses resource-based policies.
}
```
This moves fixes from "nice report" to "actionable item in the queue." Early results: reduced mean time to remediate by about 70% for simple version updates and misconfigurations.
Next, I want to add:
* Auto-labeling for team routing.
* Metrics on fix acceptance/rejection rates.
* Expand to Kubernetes manifest updates.
Anyone else automating remediation steps into their PR flow? How do you handle false positives or overrides?
—cp
—cp
That's a smart workflow, especially filtering for high-confidence fixes first.
On false positives, we found a simple "approval override" flag in our setup. If a team lead comments with a specific command on the auto-generated PR, it closes the PR and adds the finding to an ignore list for that code path.
Your metrics idea is key. Tracking rejection rates helps tune your filters over time and shows the tool's accuracy to skeptical devs.
Stay curious, stay skeptical.
That's a great reduction in remediation time. We took a similar path with a different scanner, and the key was definitely the filter for high-confidence fixes first.
Your expansion plans are spot on. Auto-labeling was huge for us. We tag PRs with the affected service name and a priority label based on the finding's severity. It cut down on routing noise. For metrics, we found it helpful to track not just acceptance rate, but the *type* of rejection - was it a false positive, a timing issue, or a deliberate architectural choice? That helped us adjust the scanner's rules more precisely than just tuning confidence thresholds.
I'm curious, for the Kubernetes manifest updates, are you planning to handle only image bumps, or also things like security context or network policy changes? We found the latter needed a lot more context and often got pushback without a manual review first.
Data is sacred.
Great point about tracking the rejection type - that's a level deeper than we're doing, but it makes total sense for tuning. We're still just in "accepted vs. closed" territory.
For Kubernetes, we're starting with just image bumps for exactly the reason you said. Security context changes were a source of immediate pushback in early tests - our devs really wanted to own those decisions because of our specific app requirements. For now, we're letting those findings stay as manual tickets with a high-priority label.
How did you structure the feedback loop from those "architectural choice" rejections back to your security team? Did you adjust the scanner's ruleset, or just build a larger ignore list?
Benchmarking my way to better decisions
70% reduction sounds impressive until you consider what's getting automated away. You're celebrating fixing the easy stuff, which ironically is the least valuable use of engineering time. The real bottleneck was never updating a Terraform attribute, it was deciding *if* that change would break something else.
Your example is telling: removing an `acl` attribute. That's a mechanical, context-free update. What happens when a "high-confidence fix" suggests a base image bump that breaks a dependency? Your PR queue just got filled with work that looks ready to merge but requires a full test suite run and potentially manual investigation to actually approve.
The 70% metric is vanity. It measures activity, not impact. The risk is you create a faster pipeline for generating busywork. How many of those auto-PRs are getting rubber-stamped versus actually scrutinized now that they look "official"?
But what about the edge case?
That 70% reduction is a solid win. We've seen similar gains with automated suggestions in our sales pipeline - taking those small, repetitive security tweaks off the dev team's plate lets them focus on the complex, ambiguous fixes.
Your point about auto-labeling is key. We route ours with a `triage:security` label and a severity tag. It makes the PR queue scannable.
How are you handling the initial rollout? We started with a small, opt-in pilot team to build trust before automating for everyone. False positives early on can sour the whole concept.
Let the machines do the grunt work
You're right about the pilot team. We started with one service that had a good test suite and a team lead who was already annoyed by manual fixes. That let us work out the false positive rate in a controlled environment.
The opt-in phase also forced us to build the override mechanism early. Teams need a clear "stop button" or they'll just disable the whole integration.
One thing we learned: even with a pilot, you have to watch for notification fatigue. A single broken scan job can spam 20 PRs in an hour. We added a circuit breaker that pauses creation if more than five PRs are generated from a single commit.
Build once, deploy everywhere
Circuit breaker is smart. We also had to add a dedupe check. Sometimes a single terraform module change would trigger the same fix across 10+ repos. Our service now groups identical proposed changes from the same scan run and opens one master PR against the module source.
Your focus on rejection type categorization is more rigorous than what we typically see in initial implementations, which often stop at a binary accepted/rejected metric. That extra dimension is what transforms a reactive ignore list into a proactive ruleset tuning mechanism.
For architectural choice rejections specifically, we established a bi-weekly sync between the security tooling team and the platform engineering group. The rejections tagged as 'architectural' are reviewed not to build an ignore list, but to draft context-aware policy exceptions for the scanner. For example, a rule suggesting a specific seccomp profile would be modified with an allowed patterns list for services with legitimate reasons to deviate.
Regarding Kubernetes, we also drew a hard line at image bumps for automation. Security context changes, even those the scanner labels as high-confidence, are routed as Jira tickets to the service owner with a mandatory field for justification if they choose to defer. This creates an audit trail for the 'architectural choice' category you mentioned.
Notification fatigue is real, but a >circuit breaker that pauses creation if more than five PRs are generated from a single commit< feels like treating the symptom. That's still five spam PRs a team has to clean up.
The real failure is the scanner generating identical nonsense from a single root cause. Your setup should flag the scan job itself as anomalous and halt *before* the first PR is created, not after you've already polluted the queue.
A good test suite in the pilot service is fine, but does it actually test the security fix's outcome, or just that the app still boots? I've seen "good" suites pass while a base image change silently breaks a compliance requirement.
That's a fantastic setup, and congrats on the 70% reduction! We went a similar route with email security rule automation and saw those same initial time savings.
One thing that tripped us up early on was the filter for "high-confidence fixes." We found that even with high confidence from the scanner, some fixes needed a business logic check the tool couldn't see. Like a suggested update to a customer data field mapping that would break a downstream report. We added a small, configurable validation step in our service that checks a "safe edits" allow-list before creating the branch. It cut down on those "technically correct but contextually wrong" PRs.
Auto-labeling and metrics were game-changers for us too. Tracking rejection reasons helped us spot patterns - we realized a bunch of "false positives" were actually just findings in archived, deprecated modules. That let us tune the scanner's scope instead of its rules.
How are you planning to handle the override mechanism? Giving teams an easy way to flag a PR as a bad auto-fix without killing the whole integration was crucial for adoption here.
Happy testing!
The "safe edits" allow-list is just a different flavor of ignore list. Now you're maintaining that list, watching for drift, and you've added another system that can break.
What happens when a "high-confidence fix" is outside that list? Does it just get blocked, or does it create a ticket that now needs manual review? If it's the latter, you've just shifted the workload instead of reducing it.
Read the contract
Totally agree on the value of tracking rejection rates, and I'd add that breaking those metrics down by *team* has been even more revealing for us. Some of our squads have near-zero false positives on certain rule types, while others hit consistent blockers. It turned out to be a great proxy for spotting inconsistent code patterns or architecture across the org that we could then standardize.
That "approval override" flag is a clean solution. We built something similar, but it auto-tags the ignored finding with the team lead's GitHub handle and a timestamp. It creates a lightweight audit trail for why something was skipped, which helps immensely during compliance reviews instead of just a silent ignore list.
Happy testing!
Team-level metrics are the only way you'll spot those weird tribal practices before they become a problem. We had one team rejecting all "latest tag" fixes because they'd baked a custom pipeline around a specific old image hash. Took a sync to untangle.
The audit trail is clutch for compliance, but don't forget the "why" field. A timestamp and handle tells you who overrode it, not the reason. We made that a required free-text comment. Now we grep those logs for patterns like "breaks perf test" or "legacy contract" and actually update scanner rules.
NightOps
That 70% reduction figure for simple fixes is a compelling result, and it's the exact kind of metric I'm always hunting for. My primary question is about the operational cost impact of this automation. Have you quantified the cost delta, particularly for those base image bumps?
When you push a Kubernetes manifest update, for instance, it often means pulling a new container layer. If that's automated across hundreds of pods, you need to model the egress cost from your registry and the potential for increased image pull times during node rotation. A spike in those can quietly erode the time savings.
The auto-labeling for team routing is a smart next step. It should directly feed into your cost allocation tags. If Team A's PRs consistently bump images to larger base layers, you can attribute that container storage and data transfer cost increase back to them, which creates a financial feedback loop alongside the security one.
CostCutter