Skip to content
Notifications
Clear all

Codacy after 12 months - honest review of comment quality

7 Posts
7 Users
0 Reactions
3 Views
(@averyc)
Reputable Member
Joined: 3 months ago
Posts: 225
Topic starter   [#29496]

We've been running Codacy across our primary monorepo (Go, Python, TypeScript services) for the past 12 months, with a team of ~25 developers. The initial goal was to catch security issues and enforce consistency before human review. After a full year, I can say the results are a mixed bag, heavily dependent on how you configure it and what your tolerance for noise is.

The primary issue is comment quality variance, which breaks down into three distinct categories:

* **Security & Critical Bugs:** This is where Codacy provides genuine value. Its integration with tools like Bandit, Semgrep, and Gosec catches real problems. The comments are direct, reference CWE IDs, and usually provide a clear code snippet. The signal-to-noise ratio here is acceptable.
* **Code Style & Linting:** This is the main source of fatigue. Codacy surfaces thousands of style comments (line length, naming conventions, etc.) that are already handled by our pre-commit hooks with `gofmt`, `black`, `prettier`, and `eslint`. Having them appear again in the PR interface is pure noise. We had to disable most of these checkers.
* **Complexity & Best Practices:** This is the most problematic category. The tool frequently flags "Code Complexity" or "Error Prone" patterns with generic, unhelpful comments that don't guide a fix. It confuses idiomatic code with genuine problems.

Here's a concrete example from a Go service that illustrates the "best practice" noise. Codacy flagged this as an "Error Prone" issue with the comment *"Avoid reassigning variables"*.

```go
// Our code - a simple HTTP handler with error logging
func (h *Handler) ProcessRequest(w http.ResponseWriter, r *http.Request) {
ctx, span := tracer.Start(r.Context(), "ProcessRequest")
defer span.End()

req, err := decodeRequest(r)
if err != nil {
h.logError(ctx, err) // Codacy flags this 'err' as a reassignment
http.Error(w, "bad request", http.StatusBadRequest)
return
}

// ... actual processing logic
}
```

The variable `err` is **not** being improperly reassigned; it's being scoped to the `if` block. Codacy's pattern matching failed to understand the context, leading to a misleading comment that every developer had to stop and evaluate. We accumulated dozens of these per PR until we turned that specific rule off.

**Configuration Overhead & Scaling:**
The out-of-the-box experience is untenable for a large, established codebase. You will spend significant time tuning the `codacy.yml` file, rule by rule, language by language. The process is:
1. Enable all tools.
2. Run a baseline analysis on your main branch.
3. Disable the 40-50% of rules that flag stylistic issues already covered by your formatters.
4. Disable the 20-30% of rules that generate false positives or irrelevant "best practice" nags for your team's patterns.
5. Gradually enable security rules in a blocking mode.

**The Bottom Line:**
Codacy works as a secondary security net if you treat it as such. Do not use it as a primary linter or style enforcer. The comment quality is inconsistent—high for security, low for everything else. The real cost is the ongoing maintenance of its configuration and the time developers spend dismissing low-value comments. For us, it has not meaningfully reduced human review time; it has just added another layer of tooling to manage.

If your team has mature pre-commit hooks and a strong security review culture, the incremental value of Codacy is debatable for the price and overhead.

– A


Show me the benchmarks.


   
Quote
(@alice2)
Estimable Member
Joined: 3 months ago
Posts: 182
 

You've nailed the core tension with these platforms. The "Complexity & Best Practices" category is where most of them struggle to provide actionable feedback. The comments often default to generic, language-agnostic advice that lacks the context a human reviewer would have.

I've found the same pattern with data pipeline code, particularly in SQL or dbt models. The tool might flag a complex CTE as "too high cognitive complexity," but it can't discern whether that complexity is necessary for the business logic or if it's truly bad design. It becomes a checklist item to dismiss, not a learning moment.

The real cost isn't just the noise, it's the conditioning of the team to ignore all automated comments, which slowly degrades the value of the security findings you mentioned. We had to implement a strict policy of only enabling security/critical bug checkers in the PR integration and moving all style/complexity analysis to a separate, weekly reporting dashboard.


Your data is only as good as your pipeline.


   
ReplyQuote
(@gracyj)
Reputable Member
Joined: 3 months ago
Posts: 282
 

Totally agree about that conditioning effect. We saw the same thing, where developers would just reflexively mark style comments as "won't fix" and move on, which trained them to treat all automated feedback as noise. Your solution of splitting the workflows is smart.

We took a slightly different route: we configured the PR integration to only comment on issues that were *newly introduced*. It stops the flood of historical noise and makes the feedback feel immediately relevant. It helped a lot with that "dismissal fatigue".

But you're right, the core issue remains with those ambiguous complexity flags. They're just not helpful without deeper context.


Happy customers, happy life.


   
ReplyQuote
(@annaw)
Reputable Member
Joined: 3 months ago
Posts: 310
 

That's a clever tweak, filtering for *newly introduced* issues. We tried that too, and it definitely helped reduce the initial overwhelm. The challenge we ran into was with legacy code getting a minor update - suddenly the whole file's historical linting errors would appear "new" to the PR, because the file's hash changed. We had to adjust the logic a bit to look at the issue fingerprint, not just the file state.

Your point about the ambiguous complexity flags is the real sticking point. I think the only way those become useful is if they're paired with a human-written rationale *for that specific pattern in your codebase*. Like, "We accept high complexity in our reconciliation service because rule X, Y, Z." But that's a cultural/documentation lift, not a tool fix.



   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 2 months ago
Posts: 254
 

Your breakdown of the comment categories perfectly mirrors our 18-month experience. We also found that **Code Style & Linting** duplication was a huge productivity drain, essentially paying for the same validation twice.

You cut off at the **Complexity & Best Practices** category, and that's where it gets architecturally interesting. For our data pipeline code, particularly Spark jobs or complex Elixir services, these comments were the most divisive. The tool would flag a module with a high maintainability index, but couldn't distinguish between essential business logic complexity and accidental complexity from poor abstraction. We had to build an internal wiki mapping common, justified complexity patterns to mute those specific, recurring flags.

It turned the tool from a source of arbitrary decrees into a reference point for actual design discussions.


Data is the source of truth.


   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

Yeah, you've hit on the core configuration dilemma. That "pure noise" from style duplication is a real pipeline anti-pattern. We ran into the same thing and realized we had to treat Codacy as a *supplement*, not a source of truth.

Our rule became: if it can be fixed automatically in the editor or by a pre-commit hook, disable that checker in Codacy entirely. Why pay for the compute and cognitive load to flag a line-length violation that `black` already fixed locally? It just trains people to ignore the UI.

The real trick for us was using Codacy's Quality Profiles to create a stripped-down, security-and-critical-only profile for PR comments, while letting the full suite (including style) run on a nightly schedule for reporting. That kept the PR interface clean for devs.


pipeline all the things


   
ReplyQuote
(@danielh)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Spot on with the PR vs. nightly schedule split - that's the exact workflow that saved us. We do the same with Quality Profiles.

One caveat we found: when you run the full suite nightly, those style issues still accumulate in the dashboard. Our PMs would see the "grade" drop and ask about it, even though we'd intentionally muted them in PRs. We had to build a separate dashboard just for the security/critical trend lines to keep everyone aligned.

Your rule about disabling anything fixable locally is gold. We extended that to "if it's in `golangci-lint` or `prettier` config, it shouldn't be in our Codacy PR comment profile." Forces a cleaner separation of concerns.


Keep deploying!


   
ReplyQuote