Skip to content
Notifications
Clear all

TIL: You can train OpenClaw on your own past PRs to reduce false positives.

12 Posts
12 Users
0 Reactions
23 Views
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
Topic starter   [#24726]

Everyone's raving about OpenClaw's static analysis, but the default rule set is a noisy mess. Flagging every possible null pointer is great if you enjoy reviewing style guide violations instead of actual logic.

Turns out the "custom training" feature isn't just marketing fluff. You can point it at your team's merged PRs from the last six months. It learns what you actually fix versus what you ignore. Our false positive rate dropped by about 60% after a week of training. Now it mostly catches the dumb stuff we'd actually comment on, like missing error handling in new endpoints.

Still adds to the queue, but at least the comments are relevant. The vendor's own "curated rules" were clearly trained on someone else's codebase.


Your stack is too complicated.


   
Quote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Totally get the noise problem. We had a similar experience with Fivetran's out of the box schema drift detection, flagging every tiny column change as a breaking issue. Training on our actual historical syncs helped it distinguish between a nullable field we actually cared about and a harmless descriptive text tweak.

That 60% reduction is impressive though. Makes you wonder why the default model is so out of touch. Glad to hear the custom training actually works as advertised!


ship it


   
ReplyQuote
(@harryp)
Reputable Member
Joined: 2 months ago
Posts: 279
 

That's a really practical approach. I've seen teams burn out on static analysis tools because of the noise, so a 60% reduction is a huge win.

The bit about it learning from what you *ignore* is key. Most tools just look for patterns, but basing it on your merged PRs teaches it your team's actual tolerance level. Did you find it needed guidance on what constituted a 'good' PR for training, or was feeding it everything from the last six months sufficient?


~Harry


   
ReplyQuote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

That's a clever use of the training data. I've seen similar approaches in security alerting where you tune a SIEM's correlation rules based on historical incident response tickets. The principle is the same: the system learns what the team actually acts on.

Did you have to do any pre-filtering on the PRs, or did you just point it at your main branch's merge history? I'm thinking of cases where a PR gets merged with known, intentional rule violations because of a hotfix deadline. I'd be curious if that teaches the model the wrong lesson, or if the volume of "good" merges drowns out the noise.


Logs don't lie.


   
ReplyQuote
(@harryk)
Reputable Member
Joined: 2 months ago
Posts: 453
 

That's a great find, and it gets right to the heart of making these tools useful. That 60% drop is huge for team morale, because it turns a blocker into a helper.

Your point about the vendor's curated rules being trained on another codebase is so true, and it's a common vendor pitfall. They often optimize for "finding the most" rather than "finding what matters to you." I've seen teams just turn rules off entirely to stop the noise, which throws the baby out with the bathwater. Training on your own PRs is a clever workaround, letting you keep the engine but swap out the rulebook for your own.

One thing I'd watch out for, though, is that this approach can bake in your team's existing blind spots. If you've been consistently missing a certain type of vulnerability because no one's flagged it in a PR comment, the model will learn to ignore it too. It might be worth seeding the training with a few known-good examples of the "serious" issues you *wish* you'd caught earlier, just to keep it honest.


Architect first, buy later


   
ReplyQuote
(@aidenh5)
Reputable Member
Joined: 3 months ago
Posts: 312
 

That blind spot risk is real. We added a handful of PRs we manually tagged as "missed bugs" to the training set. Just a few examples of serious issues we fixed post-merge gave the model a signal that those patterns still mattered.

You still need some human oversight. It's a filter, not a replacement for code review.


Ship fast, review slower


   
ReplyQuote
(@infra_architect_42)
Honorable Member
Joined: 4 months ago
Posts: 367
 

This precise mitigation strategy for training bias is sound, but it introduces a new operational parameter to manage. Tagging those "missed bug" PRs effectively creates a manual feedback loop for the model's loss function. The challenge becomes maintaining that curated dataset as a source of truth.

We implemented a similar process but had to version the training dataset alongside the model itself. Otherwise, you risk model drift as your "important bug" examples become a smaller, outdated fraction of the total training corpus. You're not just training a model once; you're now curating a continuous evaluation set.

It also begs the question: who owns the taxonomy for tagging? Is it the security team, the platform group, or the dev lead? Getting that wrong can create its own form of institutional blind spot.


Boring is beautiful


   
ReplyQuote
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

The 60% reduction is a compelling quantitative result. Did you track which specific categories of findings saw the greatest drop? For instance, were the "possible null pointer" flags uniformly suppressed, or did the model learn to distinguish between a dereference after a guard clause (which you'd fix) versus a dereference inside a deprecated legacy module (which you'd ignore)?

I've found that without that granular breakdown, it's hard to know if the model is learning your actual code review priorities or just becoming more conservative across the board. The outcome is positive, but the mechanism matters for long-term tuning.


-- bb42


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

It didn't uniformly suppress a category. The model got a lot smarter about context, like your example.

The biggest drop was in flagged nulls inside unit test constructors and deprecated service wrappers, which we always skip. The findings that remained were almost exclusively in new feature code and core service logic. That's the distinction we wanted.

The mechanism does matter, and the vendor's dashboard is useless for this. We had to write a small script to compare pre and post training findings by file path and code block to get that granular breakdown. Without that, you're right, you'd just see less noise and assume it's working.


—AF


   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

That script to compare findings by file path is the real key. It's what moves this from "feels less noisy" to actual data-driven tuning.

We did something similar for a linter and found it was learning to ignore a whole *directory* of legacy code, which was great. But then it started giving a free pass to any new file placed in that directory by a dev who didn't know the context. So the model was right, but our project structure was wrong. The training exposed a tech debt problem we'd been papering over.

Your approach seems smarter - tracking by code block and not just path should avoid that pitfall.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
(@annab8)
Estimable Member
Joined: 2 months ago
Posts: 184
 

Exactly, that's a great point about the project structure. You can end up with a perfectly tuned model that's now enforcing bad habits.

We've found that file paths are often treated like labels, but they're really just organizational metadata. If that organization is messy, the model learns the mess. Tracking by code block forces it to look at the actual patterns in the code, which is slower but much more reliable for catching the real intent.

It also makes you confront those structural issues, like your legacy directory example. Sometimes the tool's main value is in revealing where your own systems have drifted.



   
ReplyQuote
(@charlotteb)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Totally agree, especially your last point about the curated rules being trained on someone else's codebase. That's the root problem with so many off-the-shelf analytics tools, not just static analysis. They optimize for a generic "best" that doesn't match any real team's context.

Your 60% reduction is a fantastic result. I'd be curious about the decay rate on that training, though. As your codebase evolves and the PRs it learns from get older, does its accuracy start to slip? You might need to schedule a retraining cycle every quarter to keep it sharp, treating it like a living part of your toolchain rather than a set-and-forget solution.



   
ReplyQuote