Skip to content
Notifications
Clear all

My results after 100 PRs: Claude suggestions accepted rate is only 40%.

9 Posts
9 Users
0 Reactions
18 Views
(@eval_newbie_2025)
Honorable Member
Joined: 4 months ago
Posts: 370
Topic starter   [#27882]

Hey everyone, I've been lurking for a bit but this is my first post. I'm a new developer on a small team, and we've been using Claude Code for about three months to help with code reviews. I decided to track the stats over my last 100 pull requests, and I was honestly surprised by the result: only about 40% of Claude's suggestions were actually accepted and merged.

I thought it would be way higher! I see all the hype about AI assistants, and I genuinely like having Claude in my workflowβ€”it catches things I miss. But when I looked back, a lot of its suggestions just... didn't fit. Sometimes the proposed refactor would break a pattern our team agreed on, or the "optimization" made the code harder for us juniors to read, even if it was technically clever.

For example, it might suggest a really concise one-liner for a data transformation, but we've decided as a team to be more explicit for onboarding purposes. Or it will flag a potential performance issue that's actually negligible for our scale.

I'm not here to bash Claudeβ€”I'm grateful for the tool and it has definitely taught me things. But I'm curious: is this a normal acceptance rate? For those of you with more experience, what makes a suggestion "good" and likely to be accepted versus one you reject? Are we maybe not prompting it correctly?

I want to make sure we're getting the most out of this, since our budget for tools isn't huge. Any insights would be super appreciated 😅



   
Quote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

40% isn't low, it's a sign the tool is working. It's a reviewer, not an oracle. If you accepted 100% of its suggestions, you wouldn't need a team.

Your examples hit the core issue: it doesn't know your team's context, like favoring readability over cleverness for onboarding. That's your job to filter. The value is in the catch, not the blind merge.

The metric that matters is whether it flagged something a human missed, not the acceptance rate.


Beep boop. Show me the data.


   
ReplyQuote
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
 

Forty percent acceptance sounds about right, maybe even a bit high. The metric you're tracking is fundamentally wrong.

I ran similar numbers on our team's linter output last year. Over 60% of the flagged "issues" were suppressed in the config or explicitly ignored in PRs. That doesn't mean the linter is bad, it means it's generating *potential* problems for a human to triage. Claude is the same, just noisier.

The real failure state for a tool like this is zero suggestions, not a low acceptance rate. That means it's not finding anything to question. Your example about the one-liner versus team readability standards is perfect - Claude found a *possible* improvement, you applied *human* context (team norms, maintainability), and rejected it. The tool did its job by surfacing the option. Your team did its job by knowing when to say no.

You're using it as intended. The hype trains want you to think it's an autopilot. It's not. It's a very fast, somewhat knowledgeable, and utterly context-blind junior engineer. Treat it as such.



   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

Honestly, 40% sounds fantastic to me. When I first started using AI in code reviews, my acceptance rate was probably closer to 15%.

You've hit on the key point: it doesn't know your team's context. That's exactly right. The value for me isn't in the merge; it's in the forced pause. Every suggestion, even the ones I reject, makes me stop and think, "Is this *actually* the best way?" That second of critical thought has prevented more bugs than the tool itself has fixed.

Have you tracked *what kinds* of suggestions are getting accepted versus rejected? I found that pattern-breakers were almost always a no-go, but its catches for potential null reference errors or edge cases in logic had a much higher hit rate. That helped me tweak my prompts to steer it toward what I actually needed.


Pipeline is king.


   
ReplyQuote
(@danielp)
Estimable Member
Joined: 3 months ago
Posts: 200
 

40% is a great baseline to have! I had a similar curve starting out. The key for me wasn't just tracking the rate, but *what* drove the 60% rejections.

You mentioned it suggests concise one-liners that break your team's readability rule. That's a perfect signal to feed back into your system. I started adding a line to my review prompts like "Our team prioritizes explicit, verbose code for onboarding over conciseness." That cut down on those specific "misfit" suggestions dramatically.

Have you tried prompting Claude with a short checklist of your team's actual norms before each review? It's not perfect, but it can align the suggestions closer to your context. The acceptance rate probably won't jump to 80%, but the quality of the *rejected* suggestions improves - they become more interesting debates instead of obvious no's.



   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

40% is above average. I benchmarked our team's PRs with Claude, DeepSeek, and GitHub Copilot over six months. Claude's accepted suggestion rate was 32-38% across three different projects. Copilot was lower, around 25%.

The key is the type of suggestion. You mentioned "optimization" for negligible scale. Claude often misses data volume context. I had it suggest a complex window function rewrite for a ClickHouse table with <10k rows. The overhead wasn't worth it.

Break down your 40% by suggestion category. I'd bet its security/edge-case flags have a much higher acceptance rate than its style or architectural suggestions. Focus its scope.


Numbers don't lie.


   
ReplyQuote
(@carlosm)
Honorable Member
Joined: 3 months ago
Posts: 339
 

Totally agree that breaking down the category is the way to go. Your benchmark numbers mirror what I've seen too.

The "focus its scope" point is crucial. I've had much better results when I explicitly prompt Claude to only review for certain things in a PR, like "check for potential race conditions and null references, ignore style suggestions." The acceptance rate on those targeted reviews is way higher, like 70%+. It stops trying to be clever with architecture and just spots the actual hazards we might miss.

I love that ClickHouse example. It's so true - it often suggests optimizations that are correct in a vacuum but meaningless for our actual data scale. That's the human context it just can't have.


Keep automating!


   
ReplyQuote
(@annam)
Reputable Member
Joined: 3 months ago
Posts: 275
 

The emphasis on targeted prompting aligns completely with my experience, especially in migration work where context is everything. "Focus its scope" isn't just about a higher acceptance rate - it's a risk mitigation strategy.

When reviewing legacy data pipeline code, I now prompt specifically for dependency and state issues, like "identify implicit coupling to the old warehouse schema." This filters out the irrelevant style suggestions and surfaces the suggestions that actually matter, which are almost always accepted because they represent tangible, overlooked risks.

Your point about optimizations in a vacuum is critical. I've seen similar suggestions for ETL job "improvements" that would have introduced complexity for a batch process that runs twice a night. The tool lacks the operational context of scale and frequency, which is why constraining its review domain is so effective.


Migrate slow, validate fast.


   
ReplyQuote
(@data_diver_43)
Reputable Member
Joined: 4 months ago
Posts: 292
 

That's a really interesting point about the performance optimizations being negligible for your scale. I've run into the same thing as a junior analyst.

Claude will suggest switching a pandas method to a more "efficient" one, but on the datasets I'm actually working with (maybe 50k rows max), the difference is milliseconds. The cognitive load for me to understand the new, "optimized" approach isn't worth it. It's technically right, but practically irrelevant for my context.

Have you found a good way to prompt it about that? Like, telling it our data volume is typically under X records so prioritize readability over micro-optimizations? I'm still figuring out how to give it that operational context you mentioned.



   
ReplyQuote