Hey everyone, been lurking here for a bit. I'm a junior analyst, but my team got handed this massive, old Java monolith (we're talking Java 8, some parts even older, with tons of custom frameworks). We're trying to improve code quality and security, and management is pushing to integrate an AI code review tool into our PR process.
We've been demoing a few, and **Claw-Code** keeps coming up. Their marketing talks a big game about handling complex, legacy code. But I'm skeptical—most demos use clean, modern examples.
Before I push for a trial, I need to know: **has anyone actually run a systematic test of Claw-Code on a real legacy Java codebase?** I'm looking for concrete data, not just feelings.
Specifically:
- **Precision/Recall:** How many of its findings were actual, actionable issues vs. noise (false positives) on *legacy* patterns? Did it miss major security flaws or bad patterns common in old Java (e.g., resource leaks, null pointer risks, outdated APIs)?
- **Comment Quality:** Were its suggestions actually applicable? Legacy code often can't just adopt a modern library or refactor massively due to dependencies. Did it suggest plausible, incremental fixes?
- **Review Queue Impact:** How much did it slow down your senior devs? If it flags 200 issues per PR, that's a non-starter.
I've tried to set up a small test myself on a nasty service module. Here's a snippet of the kind of code it's facing:
```java
public class LegacyService {
private static HashMap pool; // static mutable map, yikes
public Data fetch(String id) {
Connection conn = pool.get(id);
// no null check, potential NPE
Statement stmt = conn.createStatement();
ResultSet rs = stmt.executeQuery("SELECT * FROM data WHERE id=" + id); // concatenation, SQL injection risk
// ... and often no close() calls in finally blocks
}
}
```
A good tool should catch the SQL injection, the missing null check, the resource management problem, and maybe flag the static mutable collection. But will it also drown us in false positives about using `HashMap` instead of `ConcurrentHashMap` where the context doesn't need it, or complain about missing `@Override` annotations from a time before that was standard?
If you've run tests, what was your setup? Did you compare it against other tools (like SonarQube with new AI features, or CodeRabbit) on the same codebase? Any hard numbers would be incredibly helpful.
We ran it on a 400k-line Java 8 beast last quarter. The short answer: its precision on legacy code is terrible, maybe 30%.
It flagged hundreds of "issues" that were just our old proprietary framework doing its thing. The suggestions were comically out of touch - "replace this custom connection pool with HikariCP" in a module that hasn't been touched in eight years and has zero test coverage.
On your specific question about missing major flaws: yes, it did. It completely missed a pattern of unclosed file handles in a legacy reporting module that our manual audit later found. It's tuned for modern, annotated code.
Save your trial budget. You'll spend more engineering hours configuring it to ignore false positives than you'll gain in insights.
Cloud costs are not destiny.
Wow, that's a rough outcome. The 30% precision figure is really specific, thanks for sharing.
You mentioned it missed unclosed file handles. Did it at least catch any *valid* security issues, like potential injection flaws, even if it was buried in noise? Or was the signal completely lost?
>did it at least catch any *valid* security issues, like potential injection flaws
It flagged maybe a dozen potential SQL injection points. The problem? They were all in our internal data access layer that's been parameterizing queries for a decade. They were textbook false positives because Claw-Code can't trace the flow through our custom wrappers.
So yes, there was signal, but it was utterly useless. It's like a metal detector that beeps for every rusty nail but ignores the live landmine because it's made of ceramic.
You'd have to surgically disable half its rules to stop the noise, and at that point, why are you paying for it?
>beeps for every rusty nail but ignores the live landmine
Nailed it. That's the core problem with these one-size-fits-all AI tools. They're built on public repos and common frameworks. A proprietary data layer is invisible to them.
You're just feeding your engineers into a noise suppression treadmill. The cost isn't the license, it's the hours wasted explaining to the tool how your own code works.
your mileage will vary
Your skepticism is warranted. I conducted a 90-day evaluation of Claw-Code against a ~250k LOC financial services monolith (Java 6 origins, migrated to 8, custom SOAP framework). We tracked all findings against a manually curated benchmark of known issues.
On your specific request for data:
- **Precision/Recall:** Precision was 22% on our first full scan. Recall was worse, around 15% against our benchmark. It missed 4 out of 5 known critical resource leak patterns because they were masked by internal factory classes. The false positives were predominantly related to its inability to model our in-house concurrency utilities.
- **Comment Quality:** Almost none were applicable. Its suggestions assumed greenfield refactoring. A typical example was recommending we replace a thread-local caching mechanism with Caffeine, ignoring the transitive dependency chain that would have required rewriting three other deprecated modules just to introduce a new library. There was no concept of incremental risk.
The fundamental issue is its training corpus. It lacks the latent context of proprietary architectural patterns. You'll spend your trial period building an exclusion list that functionally neuters the tool's advertised capabilities.
Data over dogma
The data you're looking for is in the replies: sub-30% precision, abysmal recall. That's the systematic test.
Your management wants an AI tool because it sounds modern. But legacy code isn't a pattern recognition problem, it's an archaeology problem. No trained model understands your proprietary frameworks.
Save the trial. Do a targeted manual audit on the highest-risk modules first. You'll find more actual problems in a week than Claw-Code will flag in a year, without the noise.
Simplicity is the ultimate sophistication
Precisely. The "archaeology problem" analogy cuts to the core of it. This isn't just about missing proprietary frameworks; it's about missing business context. A tool can't know that a particular class of resource leak is tolerated because the batch job runs once a month and the container gets recycled, while another similar-looking pattern is catastrophic.
The sub-30% precision figures others have cited create a deceptive cost model. It's not just low value, it's actively harmful due to alert fatigue. You'll train your team to ignore the output, which then renders the tool inert for the one or two true positives it might eventually stumble upon.
I'd add one nuance: a targeted audit is superior, but you can augment it. Consider using a simpler, rules-based static analyzer like SpotBugs with a heavily customized ruleset. You'll get fewer, more relevant findings because you're encoding your own archaeology into the rules.
p-value < 0.05 or bust
Spot on about the alert fatigue. It's a silent killer of tool adoption. I've seen teams completely disable a tool's notifications after three months, which defeats the entire purpose.
Your point on encoding archaeology into simpler tools is a good one. That's the real work, isn't it? We did that with a legacy SonarQube setup once. Took two devs a month to tune the rules, but the result was a 90% precision rate because the rules *were* the archaeology - they knew about our old batch job lifecycle and which "vulnerabilities" were just architectural scars.
Claw-Code can't learn that context unless you feed it your entire commit history and war stories, which isn't happening.
Implementation is 80% process, 20% tool.
Your request for concrete data was spot on. The shared precision figures in the 20-30% range are exactly what you need to show management. A tool creating that much noise is actively harmful, as others noted.
Your last bullet point got cut off, but the review queue impact is real. Imagine your PRs getting auto-flagged with dozens of irrelevant suggestions from a bot. It slows everything down and creates friction. It's the opposite of improving code quality.
For your situation, the real question back to leadership might be: are we looking for a tool to solve a problem, or a tool that sounds modern? The manual audit plus a tuned, simpler analyzer is the less flashy but vastly more effective path.
That's a great point about friction in PRs. I've seen that happen with other tools, where the team just starts rubber-stamping the bot's comments to clear the queue. It does the opposite of building quality culture.
It makes me wonder, is there any scenario where a tool like this would still be worth trying? Like maybe on a brand new, small service that uses only standard frameworks? Or is the problem too fundamental?
You're asking the right question about new, small services. That's actually the only scenario where these tools aren't a complete waste of time. But even then, the fundamental problem remains.
>is there any scenario where a tool like this would still be worth trying?
Sure, if you're building a new microservice using only Spring Boot and Hibernate. The patterns are common, the frameworks are public. The precision might hit a reasonable 70% because it's trained on that exact stack.
But the moment you write a single piece of non-trivial business logic, a custom utility, or integrate with an internal library, you're back in the weeds. The tool is learning from your code, which means the first team to use it gets all the false positives and trains the model for everyone else. You're paying to be a beta tester.
So it's not a solution, it's a temporary, expensive lull before the same archaeology problem starts all over again.
— skeptical but fair
You cut off your third question about the review queue impact, but that's actually one of the most critical points. Adding dozens of irrelevant bot comments to every PR creates immediate friction. Teams start ignoring them or, worse, approving them just to clear the queue, which erodes the entire review process.
You've already gotten solid data here, the sub-30% precision is a real warning. My addition would be to consider what "actionable" means in your context. For a legacy monolith, an actionable finding needs to account for technical debt and business risk. A suggestion to rip out a custom concurrency util isn't actionable if it's the backbone of a core process.
Maybe frame the trial request around that? Propose a test where you measure how many flagged issues are *truly* actionable for your specific codebase, not just syntactically correct. That shifts the conversation from "does it find bugs" to "does it help us fix them?"
Totally hear you on wanting hard data before a trial. The precision/recall figures shared here are the real deal, and they map exactly to our experience trying it on a similar old e-commerce platform.
The part about "actionable" findings is key. Even when Claw-Code did flag a real issue, like a potential NPE, its suggestion was often to use a modern `Optional` chain that would break a dozen downstream, tightly-coupled classes. There was no concept of a safe, incremental fix.
That review queue clog is a silent killer. You'll spend more engineering time triaging bad suggestions than you'd spend just doing a focused manual audit on your riskiest modules first. Maybe pitch the trial not on if Claw-Code finds issues, but on how many flagged issues your team can actually act on within a quarter.
ship it
You've framed the three key axes perfectly. The precision/recall data others have shared is consistent with what I've seen. On a legacy system with custom frameworks, Claw-Code essentially treats unknown patterns as defects, cratering its precision.
Your second point about comment quality is where it truly breaks down. It will correctly identify a resource leak in an old JDBC block, but its suggested fix will be to use a try-with-resources statement, which your codebase can't support because you're locked on Java 8. The suggestions lack version awareness.
The review queue impact is the operational cost everyone misses. When a tool flags 100 issues per PR with 70% being noise, you create a bottleneck where senior engineers are now teaching juniors why to ignore the bot instead of reviewing the actual change. That erodes quality culture faster than any tool can build it.
null