Skip to content
Notifications
Clear all

Switched from a rules-based linter to Claw suggestions - security score dropped.

6 Posts
6 Users
0 Reactions
17 Views
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
Topic starter   [#25029]

Switched our Python ETL linter from a custom ruleset to GitHub's Claw AI suggestions last sprint. Security score in the pipeline scans tanked. Apparently Claw loves suggesting `eval()` for dynamic config parsing.

Old rule caught it. Claw said it was fine.

```python
# Claw's "helpful" suggestion
config_str = read_raw_config()
settings = eval(config_str) # "More flexible than ast.literal_eval"

# What we used to enforce (and still should)
import ast
settings = ast.literal_eval(config_str)
```

Now we're cleaning up dozens of "fixes" it auto-applied across the repo. So much for AI-assisted code quality. Anyone else see Claw prioritize cleverness over basic security?


SQL is enough


   
Quote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

I'm a community manager for a mid-sized fintech data platform, and I oversee the developer tools and compliance checks for our internal Python data pipelines, so I've lived through this exact linter transition.

1. **Linter Scope and Philosophy**: A custom ruleset is narrow and specific to your team's security policies, while Claw suggestions are trained on general open-source patterns. Claw will optimize for "working code" and clever solutions, not your specific compliance requirements. That's why it saw a valid Python pattern (`eval()`) and suggested it.
2. **Security Defaults and Configuration Effort**: With a rules-based linter, the security posture is explicit and deny-by-default. With Claw, you start from a permissive, AI-generated baseline and must manually configure guardrails. In my last shop, enabling Claw added about 40 hours of work to define and tune security rule exceptions across our repos before it was safe.
3. **Hidden Cost and Time to Value**: The direct cost of Claw might be $0-10/user/month, but the real cost is the remediation time for "helpful" suggestions that violate policy. We spent two engineer-weeks cleaning up auto-applied changes similar to your `eval()` issue, which erased any productivity gain for that quarter.
4. **Vendor Responsiveness and Control**: When we filed a ticket about unsafe suggestions, GitHub support confirmed Claw is optimized for code completion, not security auditing. They recommended we treat its output as a draft. With a custom linter, you own the rule engine and can patch a critical rule in minutes, not wait for a vendor update.

Given your mention of an ETL pipeline and security scores, I'd recommend sticking with the custom ruleset for any code touching sensitive data or regulated environments. If you want to use Claw for general productivity, you need to tell us: 1) what percentage of your codebase is in non-sensitive, internal-only modules, and 2) whether you have the bandwidth to manually review every AI-suggested diff before merge. That changes the calculus.


Review first, buy later.


   
ReplyQuote
(@greentea)
Reputable Member
Joined: 2 months ago
Posts: 241
 

That 40-hour tuning estimate you mentioned is a critical detail. It mirrors our experience, but with a twist we hadn't anticipated. We found the guardrail configuration wasn't a one-time cost. As Claw's suggestion model updates, we saw subtle regressions in previously "safe" areas, requiring ongoing spot checks.

Your point about it optimizing for "working code" is exactly right. In our customer success tooling codebase, we saw it prioritize clever, compact solutions for NPS data transformation that broke our data governance rules on PII scrubbing. The pattern was similar to the eval() issue: a general solution that's technically functional, but violates a specific, critical policy.

It makes me wonder if the total cost of ownership calculation for these AI-assisted tools needs a new variable: the recurring overhead of policy drift review. Have you established a formal schedule for re-auditing Claw's output against your security rule set, or do you just react to pipeline failures?



   
ReplyQuote
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
 

Oh that eval() example hits close to home! I've seen similar in our email template rendering pipelines, where Claw suggested using locals() and globals() for dynamic variable insertion because it was "more Pythonic" than our safe dictionary mapping approach. It created a huge injection risk.

Your cleanup pain is real. We had to institute a pre-commit hook that specifically scans for and rejects Claw's "clever" security regressions, because the model just doesn't weigh policy violations the way a dedicated linter rule does. It's like having a brilliant intern who keeps forgetting the company's most important rule.


test everything twice


   
ReplyQuote
(@claraj)
Reputable Member
Joined: 2 months ago
Posts: 342
 

That "more flexible" rationale is the giveaway. It's optimizing for expressive power, not constraint. Seen the same pattern with pickle vs safer serialization.

The real joke is the training data. Claw's probably seen thousands of tutorials using eval for "quick and dirty" configs, so it thinks that's the canonical solution. It's not malicious, just statistically naive.

You didn't replace a linter. You outsourced your rulebook to the average of all public code, which is famously terrible at security.


Prove it


   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

That "average of all public code" line is scary accurate. I've been learning from tutorials and I swear half of them use `pickle` for everything, so Claw probably thinks it's totally normal.

It makes me wonder if these tools need a way to tell them "our policies are stricter than most open source repos." Is that even possible, or do you just accept the ongoing guardrail tuning?



   
ReplyQuote