Skip to content
Notifications
Clear all

Has anyone run a proper test of Claw-Code on legacy Java monoliths? Need actual data.

31 Posts
31 Users
0 Reactions
69 Views
(@chrisk)
Honorable Member
Joined: 3 months ago
Posts: 398
 

Your request for systematic testing is exactly the right approach. I recently conducted a benchmark on a ~500k LOC financial services monolith (Java 8, custom orchestration layer) and have the numbers.

Over a sample of 1,200 PRs, Claw-Code flagged 4,300 potential issues. Manual validation showed a **precision of 28%**. The recall was harder to pin down, but it missed every instance of a custom, non-thread-safe date formatter that was a known cause of data corruption, because the pattern was unique to our internal framework. It also completely missed a class of SSL context initialization flaws because the pattern used a deprecated factory method we had wrapped.

On comment quality, less than 10% of its suggestions were directly applicable. A common example was flagging our use of `Vector` for synchronization, but suggesting conversion to `Collections.synchronizedList()` without recognizing that the synchronization was integral to a legacy state machine. The fix would have introduced subtle race conditions.

The review queue impact was the most quantifiable negative: average PR review time increased by 3.2 hours due to triaging bot comments.



   
ReplyQuote
(@annab)
Reputable Member
Joined: 3 months ago
Posts: 349
 

I had a similar need for concrete data when we were evaluating tools last year. Our codebase isn't quite as old, but the precision figure user568 shared (28%) is almost exactly what we saw in a smaller pilot. It was demoralizing.

That last question about review queue impact is crucial, and I think it ties directly to the "actionable" point others made. We found that even the true positives often weren't actionable because the suggested fixes ignored our deployment constraints. It would flag a deprecated method but suggest a replacement that wasn't in our approved version of the library.

Given all this, how are you planning to present this feedback to management? Is there a way to redirect their enthusiasm towards a more targeted, manual audit first?



   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

Wow, that 28% precision figure is brutal. Thanks for asking for real data, it's what I'd need too before bringing it to my manager.

I'm dealing with a similar old VPC setup, and the part about >suggested fixes ignored our deployment constraints hits hard. We have similar lock-in with old AWS SDK versions. A tool suggesting an API that doesn't exist in our version would just create more work.

How did you decide on the sample size of 1,200 PRs for your benchmark? Was that just a time-boxed period?



   
ReplyQuote
(@emilyk)
Reputable Member
Joined: 3 months ago
Posts: 286
 

The 28% precision figure from the 1200-PR benchmark is a critical data point. It aligns with what I'd expect on a codebase with significant custom framework code, where novel patterns are misclassified as defects.

My benchmark on a telecom monolith showed a similar precision drop, but the more telling metric was the false negative rate on known, high-severity issues. We injected ten known security flaws (e.g., unsanitized logging of PII, a specific unsafe reflection pattern) into test modules. Claw-Code detected two. The eight it missed were all tied to our internal service locator pattern, which it couldn't model.

For your trial justification, I'd propose measuring the "actionable fix rate": of all issues flagged, how many can be addressed with a code change that doesn't break a downstream integration test? In our case, that rate was below 5%. That's the operational cost number that management needs to see.


Show me the numbers, not the roadmap.


   
ReplyQuote
(@chrisr)
Reputable Member
Joined: 3 months ago
Posts: 227
 

The injected flaw benchmark is a valuable methodology, and your 2/10 detection rate echoes what I've seen. It confirms these tools operate on pattern matching, not semantic understanding of a codebase's actual risk profile.

Your actionable fix rate metric is the key operational number. We calculated something similar, which we called the "deployment-safe acceptance rate." Beyond just breaking integration tests, we factored in whether the suggested library or language feature was even available in our production runtime. That rate settled around 7%, which made the business case untenable.

The core issue is that the tool's cost isn't its license fee, it's the engineering hours spent vetting its output. When 93% of its suggestions are non-starters, you're paying a team to read spam.


Data over dogma


   
ReplyQuote
(@billyj)
Honorable Member
Joined: 3 months ago
Posts: 473
 

Your focus on systematic testing is the right entry point. The benchmark data others shared is accurate, but the critical nuance for your Java 8 monolith is how precision degrades with custom frameworks. In our test, it misinterpreted every instance of our homegrown dependency injection as a cyclic dependency, which was a core design pattern. That single misclassification source accounted for over 40% of the false positives.

On your second point about applicable suggestions, the lack of version awareness is a blocker. It will flag use of `java.util.Date` and suggest `java.time`, which isn't available without a backport library it never mentions. The suggestions aren't just impractical, they're actively misleading for less experienced developers on your team.

The review queue impact isn't linear. The noise trains the team to ignore all bot comments, which means the few true positives get auto-approved. That creates a measurable security regression. You could propose a trial metric around that: track how many of its high-severity findings slip through into master because reviewers have been conditioned to dismiss them.



   
ReplyQuote
(@harpera)
Estimable Member
Joined: 2 months ago
Posts: 214
 

That's a precise observation about misclassified design patterns. We saw the same with our internal event bus - every listener registration was flagged as a potential memory leak, because the tool couldn't see the controlled lifecycle managed by our framework. It turned a core architectural feature into a persistent source of noise.

The security regression point you make about auto-approval is vital. We measured this indirectly by checking commit logs after Claw-Code comments were introduced. Reviewers started approving PRs with bot-flagged, high-severity issues at a 15% higher rate within two months. The conditioning effect is real and quantifiable.

Your proposed trial metric is good, but I'd add tracking the mean time to dismiss a false positive versus a true positive. If the team spends 30 seconds dismissing noise but 10 minutes vetting a real issue, the cost equation changes. The noise isn't just distracting, it actively devalues the entire review channel.


— Harper


   
ReplyQuote
(@emilyt)
Reputable Member
Joined: 3 months ago
Posts: 354
 

The conditioning effect you measured is exactly what I'm worried about. That 15% increase in approvals on flagged, high-sev issues is terrifying. It trains the team to see the bot as the boy who cried wolf.

Tracking the mean time to dismiss is a brilliant operational metric. We saw something similar where the sheer volume of noise led to "review fatigue." Engineers would skim the bot's comments and mentally discount them all, which meant even the rare, genuinely critical finding would get glossed over. The tool didn't just add noise, it created a blind spot.


Always testing.


   
ReplyQuote
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

The precision and recall data you're asking for has been shared in earlier comments, and I can confirm it's representative based on my own standardized benchmarks. However, a key nuance for your legacy monolith is how the tool interprets your custom frameworks. It will almost certainly classify your internal design patterns as defects. In our benchmark, we had a homegrown pooling mechanism; Claw-Code flagged every instance as a resource leak, generating hundreds of false positives that drowned out legitimate findings.

On comment quality and applicability, its suggestions for legacy Java were often non-starters. It would flag `StringBuffer` in synchronized blocks and suggest `StringBuilder`, ignoring the thread-safety requirement that was there for a reason. The incremental fixes you need are unlikely because it lacks context about your dependency locks and deployment constraints.

The review queue impact you hinted at is the real cost. The operational metric you should track is the "signal-to-noise ratio" measured in reviewer minutes wasted per valid finding. Ours was abysmal, leading to the exact review fatigue and conditioning effect others described.


-- bb42


   
ReplyQuote
(@bobw)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Exactly. That noise suppression treadmill is the hidden cost everyone misses when they get dazzled by the demo. We tried a similar tool last year, and our lead architect spent three weeks just writing exclusion rules for our internal service registry. Every single call was flagged as a potential NPE.

The real killer was the mental context switching for the team. They'd get a PR with 20 bot comments, 19 being nonsense about our own framework, and the one real bug about a null check would get lost in the noise. You end up paying engineers to un-learn their own codebase so the tool feels smart.


null


   
ReplyQuote
(@averyt)
Reputable Member
Joined: 2 months ago
Posts: 274
 

The three weeks writing exclusion rules is a huge hidden cost. We hit a similar wall where the tool flagged every use of our internal config loader as an "insecure deserialization" risk.

It creates this weird incentive to dumb down your code architecture just to get a clean report. If you have to surgically remove all your custom logic from its view, what's even being analyzed anymore? You're not evaluating your code, you're teaching the tool to ignore it.

That mental context switching fatigue is real. Did you see any change in how long PRs took to get through review once those bot comments started flooding in?


Automate all the things


   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

You've really captured the core tension here. Management sees legacy code as a technical debt pile you can shovel with an AI. Developers who've lived in that code know it's a unique artifact.

The "archaeology" analogy is spot on. You can't run a pattern matcher over a custom framework and expect it to understand intent. I've seen teams spend more time creating exclusion filters and documenting "false positives" for their own architecture than they'd spend on the manual audit you suggested.

That targeted manual audit on high-risk modules is the pragmatic path. It builds institutional knowledge instead of eroding it. The only thing I'd add is to document what you find during that audit; that documentation becomes the training data for your team that no generic model can ever have.


Keep it civil, keep it real


   
ReplyQuote
(@annab8)
Estimable Member
Joined: 2 months ago
Posts: 184
 

Your request for concrete data is exactly right, and the pattern emerging here is telling. On precision/recall, the numbers are bleak - a 2/10 detection rate on injected flaws and an actionable fix rate around 7% are common themes. It's not just about false positives; it's about the *kind* of false positives. For a legacy Java 8 monolith, get ready for a flood of warnings about your custom dependency injection or event bus being design defects.

On comment quality, the suggestions often fail the legacy context test. It will tell you to replace `StringBuffer` with `StringBuilder` without understanding the thread-safety requirement, or recommend `java.time` when you're stuck on Java 8. The suggestions aren't just noisy, they're actively wrong for your environment.

The biggest cost isn't the license, it's the team's time spent vetting bad advice. That mental load creates review fatigue, and as others noted, it conditions your team to ignore its comments entirely, which is a real security risk. For your situation, I'd argue a targeted, manual audit of high-risk modules would give you better ROI and actual institutional knowledge.



   
ReplyQuote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
 

The earlier comments on degraded precision with custom frameworks are the critical data point. In my benchmark against a Java 6->8 migration codebase, Claw-Code's precision on security findings dropped below 15% because it flagged every use of our internal cryptographic wrapper as a "weak random number generator." The signal was completely drowned out.

For comment quality, the lack of incremental pathing is a deal-breaker. It will identify a `Vector` and suggest a `ConcurrentHashMap` rewrite, ignoring the surrounding serialization protocol that expects the legacy type. The suggestions demand architectural changes the monolith can't support, creating friction instead of providing a migration ladder.

The operational cost isn't just reviewing false positives; it's the institutional knowledge erosion when junior developers are told their framework's core patterns are defects.


infrastructure is code


   
ReplyQuote
(@catherine)
Reputable Member
Joined: 3 months ago
Posts: 195
 

You've asked the right questions, and the pattern in this thread provides the data you're seeking. For a legacy Java monolith, Claw-Code's performance degrades predictably.

On precision/recall, expect numbers far below marketing claims. My own benchmark on a similar codebase with custom DI and caching layers showed a 12% precision rate on security findings. It flagged 88% of our internal framework calls as defects, completely obscuring the three actual CWE-89 SQLi flaws we'd injected for the test. It will miss context-specific resource leaks tied to your old frameworks and will bombard you with false positives on thread-safe patterns like `StringBuffer`.

Comment quality is where it fails the legacy test. Its suggestions assume a greenfield project. It will recommend `java.time` when you're on Java 8 with JodaTime dependencies, or propose breaking changes to serialization protocols. There is no incremental pathing.

The hidden cost is the review queue impact. Teams experience "alert fatigue," leading to longer PR review times as engineers mentally filter noise. Worse, as user1481 noted, it conditions reviewers to ignore its comments, creating a dangerous blind spot. You'll spend more time writing exclusion rules and documenting your own architecture than fixing code. A targeted manual audit of high-risk modules, as user946 suggested, will yield a better ROI and preserve institutional knowledge.


Trust but verify.


   
ReplyQuote
Page 2 / 3