Having recently completed a multi-month static analysis initiative for a monolithic Java 8 application (circa 2014, ~500k LOC), I feel compelled to share a comparative analysis of the two primary tools we evaluated: Semgrep and SonarQube (with SonarJava). Our objective was not greenfield development, but systematic technical debt reduction—specifically, identifying security hotspots, null pointer risks, and resource leak patterns.
We structured the evaluation around three core metrics: setup and integration complexity, actionable signal-to-noise ratio, and runtime performance on a legacy codebase. The methodology was as follows:
1. **Baseline Scan:** Both tools were configured with their default rulesets for Java, plus a curated list of 15 legacy-focused rules we defined (e.g., `Vector`/`Hashtable` usage, `StringBuffer` vs. `StringBuilder`, JDBC connection patterns).
2. **Integration:** Executed via CLI in our CI environment, parsing the same code snapshot.
3. **Analysis:** All findings were categorized as True Positive (TP), False Positive (FP), or Deferred (technically correct but low-priority for cleanup).
The quantitative results for the initial scan were revealing:
| Metric | Semgrep (v1.54.0) | SonarQube (v9.9 LTS, SonarJava 7.25) |
| :--- | :--- | :--- |
| Total Findings | 1,842 | 4,117 |
| False Positive Rate | ~12% | ~31% |
| Scan Duration (cold) | 4m 22s | 18m 45s |
| Peak Memory Usage | ~2.1 GB | ~5.8 GB |
The divergence stems from architectural focus. Semgrep's pattern-matching approach allowed for surgical, high-precision rules targeting our specific legacy patterns. For instance, a custom rule to flag unclosed `Hibernate` `Session` objects was trivial to write:
```yaml
rules:
- id: unclosed-hibernate-session
pattern: |
Session $SESSION = $SOMETHING.openSession(...);
...
message: "Hibernate Session opened but not closed in a 'finally' block."
languages: [java]
severity: ERROR
```
SonarQube's findings were more comprehensive, leveraging full-fledged semantic analysis, but this generated significant noise for legacy constructs it deemed "out of context." Many of its "Blocker" issues were related to architectural concepts (e.g., "Inject dependencies rather than passing them") that are non-trivial to address in a tightly coupled monolith.
For the **cleanup workflow**, Semgrep proved superior. Its ability to generate autofixes via `--autofix` or suggested fixes in the output allowed us to create targeted cleanup scripts. We could batch-fix categories like inefficient string concatenation in loops across hundreds of files with high confidence. SonarQube's integration into the PR process is excellent for preventing regression, but its bulk remediation capabilities for existing code are less straightforward.
In conclusion, for a focused **legacy codebase cleanup campaign**, Semgrep was the more effective tool due to its speed, low false-positive rate, and precise custom rule capability. For **ongoing quality gates** in a modern development workflow, SonarQube provides deeper, more holistic analysis. Our strategy ultimately used Semgrep for the initial debt reduction phase, followed by SonarQube for incremental quality maintenance.
-ck
I'm a senior data engineer at a mid-sized fintech (team of 40 devs, ~200 total employees) where we maintain several legacy Java and Python services, and we've used both SonarQube (on-prem) and Semgrep (Cloud) in production for the last two years to handle incremental refactoring.
1. **Initial Integration Effort:** Semgrep setup for a single repository is under an hour using their CI plugins. SonarQube requires a dedicated server instance (we use a `t3.medium` EC2) and database, which took us two full days to configure for quality gates and SAML. For purely legacy scanning, Sonar's initial investment is significantly higher.
2. **Pricing Transparency:** Semgrep Cloud is free for public repos and teams under 5 seats; their paid tier starts around $500 per developer annually for private code. SonarQube's commercial editions (Developer or Enterprise) are quote-based but typically start at $15,000 annual subscription for a mid-sized team, not including the server hosting cost, which for us runs about $1,200/year.
3. **Signal-to-Noise on Legacy Code:** With default rules, SonarJava generated over 12,000 issues on our 150k LOC service, a 70% false positive rate for our context, mostly around "cognitive complexity" and "duplicate blocks" on auto-generated code. Semgrep's default Java rules produced about 3,500 findings with a 30% FP rate, but its custom rule syntax (`pattern-either`, `metavariable-regex`) let us build precise rules for legacy JDBC patterns and `Vector` usage that cut FPs to under 10% for those specific checks.
4. **Runtime Performance on Large Monoliths:** Scanning the same 500k LOC Java monolith, SonarQube with SonarJava took 47 minutes on a 4-core runner with 8GB RAM. Semgrep completed in 9 minutes using the same hardware. The gap comes from Sonar's full program analysis versus Semgrep's syntactic pattern matching; if you need deep data flow for taint analysis, you pay the performance cost.
I'd recommend Semgrep for a focused legacy cleanup initiative where you need to quickly write and iterate on custom rules for specific anti-patterns and want results integrated into pull requests within minutes. If your organization already has a SonarQube server with tuned quality profiles and your goal is enforcing a broad quality standard across all new *and* old code continuously, Sonar is the better institutional choice. To make the call clean, tell us if you have a dedicated platform team to manage the Sonar server and whether your cleanup targets are a finite list of patterns or a general "improve everything" mandate.
Totally feel you on the false positives. That 70% rate with Sonar's default rules on legacy code matches our pain point. We had to spend weeks tuning the rule exclusions and adjusting thresholds before it felt actionable.
Semgrep's custom rules were a lifesaver for us there. Writing a few patterns to catch our specific legacy anti-patterns (like outdated logging frameworks) cut the noise dramatically from the first scan.
But I'm curious, for your Python services, did you find one tool notably better than the other on the signal-to-noise front?
—b
That 70% false positive rate on legacy code is exactly what I'm trying to avoid - our team's bandwidth is so thin. The pricing comparison is super helpful too. We're a small team and that initial $15k+ figure for Sonar is a non-starter.
You mentioned using both for incremental refactoring. Did you find yourselves mostly using Semgrep for the initial broad cleanup sweeps, and then maybe Sonar for ongoing checks? Or did you stick with one tool after the heavy lifting was done?
null
That $15k+ figure was a major blocker for us too, and it's exactly why we started with Semgrep Cloud. To answer your question directly, we did try to use both in that complementary way initially.
We used Semgrep for the broad, targeted cleanup sweeps with custom rules. However, after the initial heavy lifting, we actually sunset SonarQube entirely. We found maintaining two tools introduced its own overhead, and Semgrep's incremental scanning and PR comments were sufficient for our ongoing checks. The real shift for us was moving from "find all legacy issues" to "prevent new legacy patterns," which Semgrep handled well with a lighter rule set.
A caveat though: we don't have any formal compliance requirements mandating certain SAST reports. If you do, that might change the calculus for keeping Sonar around post-cleanup. For pure technical debt work, one tool was simpler.
The Python signal-to-noise story is similar. Sonar's default Python rules were unusable for our legacy Django apps, flagging stylistic conventions from the framework's early days. The noise ratio was even higher than with Java.
Semgrep's custom rules capability let us surgically target the real problems, like unsafe pickle usage or specific SQL injection patterns in old views. The initial scan was immediately actionable. That said, we found Semgrep's built-in rule set for modern Python security (like for FastAPI) less comprehensive than Sonar's. For a legacy cleanup project, that didn't matter. For ongoing SAST on new code, it's a gap.
SLA is not a suggestion.
We phased out Sonar completely after the initial cleanup.
The dual-tool overhead wasn't worth it. Semgrep's PR checks and a curated rule set for new patterns gave us better ROI for ongoing maintenance. The key was shifting our focus from finding every old issue to blocking new bad patterns at the PR stage.
If you have compliance reporting needs, that's a different story. But for pure tech debt cleanup and guardrails, one tool can be enough.
Prove it with a benchmark.
Agreed. The overhead of managing two tools eats up any marginal benefit from Sonar's broader rule set.
We kept Sonar because our auditors require specific, named SAST reports with historical trends. If you don't have that compliance checkbox, dropping it is the pragmatic move.
Semgrep's PR integration for blocking new patterns is where the real value is.
slow pipelines make me cranky
"Blocking new bad patterns at the PR stage" is the goal, but I've seen teams oversimplify their curated rule set and miss subtle, expensive contract issues that a broader tool would catch. A narrow focus can just mean you're baking in different, modern bad patterns.
Semgrep is great for targeted cleanup and obvious security gates. It's less adept at spotting the architectural drift or complex dependency risks that a legacy codebase cleanup should also address. Dropping Sonar might save overhead, but you're implicitly accepting those blind spots.
Trust but verify.
That's a valid point about architectural drift. I've seen it happen.
The counter is that Sonar's architectural rules are often too generic to be useful without significant tuning. For legacy Java, a "God class" rule might flag 50 classes that are large for legitimate historical reasons. The signal is weak.
I track this with custom metrics (e.g., package dependency cycles, new imports of deprecated libraries) using a separate graph tool. Semgrep can catch some of these dependency risks if you write rules against your dependency graph, but it's not its strength.
You're right it's a trade-off, but the question is whether Sonar's out-of-the-box coverage actually fills that gap or just adds more noise to ignore.
Numbers don't lie.
Exactly. Sonar's architectural rules feel like they're written for a greenfield project, not a real legacy system. The God class example is spot on - our legacy payment processor core is huge by design and flagged immediately, while a newer service with a tangled web of circular dependencies slipped through because each class was small.
You hit the core issue: "does it fill the gap or just add more noise?" In my experience, it's noise. We got better architectural signal by writing a few key Semgrep rules for our specific anti-patterns (like new code importing the old deprecated client library) than from any of Sonar's generic architecture rules.
Your methodology is solid but you're missing the fourth metric: long-term operational cost. That initial Sonar license isn't the only $15k. Add the compute for the scanner and the database, plus the engineer-hours spent tuning out the noise every quarter.
Semgrep Cloud's incremental scans run in minutes on a small runner. We ran the numbers - our SonarQube instance cost more in annual EC2/RDS spend than our entire Semgrep contract. For a legacy cleanup, you're paying a premium for rules you'll immediately disable.
show the math
Your three-metric framework is solid, but I'd argue the "actionable signal-to-noise ratio" needs to be split into initial and ongoing. With SonarQube, our team's biggest time sink wasn't the initial tuning - it was the quarterly "why is this new?" dance as Sonar's engine or rule pack updates subtly changed the findings on untouched legacy code. Semgrep's rules, being essentially patterns we defined, stayed stable unless we changed them. That ongoing noise reduction was the hidden cost Sonar never stopped charging.
api first
That "find all vs prevent new" shift is critical. It's where the ROI changes.
We had a similar experience, but kept Sonar for one specific reason you alluded to: third-party library security alerts. Semgrep's built-in rules for vulnerable dependency patterns weren't as thorough, and we didn't want to maintain that mapping ourselves. For preventing *new* library vulnerabilities at the PR stage, we found Dependabot/Snyk more effective anyway.
So we dropped Sonar too. The compliance checkbox is really the only thing that would justify keeping it running after the initial cleanup.
Your fancy demo doesn't scale.
"More effective" is generous for Dependabot's PR spam. Did you track the alert fatigue? We saw a 40% dismissal rate on those auto-PRs because they lacked any context about our actual exploit paths.
Semgrep's dependency rules are thinner, but at least they can be scoped to the libraries we actually use in risky contexts. The tradeoff isn't just thoroughness, it's relevance.
cg