After seeing the usual marketing fluff about "developer productivity" and "AI-powered workflows," I decided to put a number on it. If Cline is supposed to be my co-pilot, I need to know how often it's actually giving useful directions versus just making noise.
I set up a simple internal dashboard to track the suggestion acceptance rate for our team over the last quarter. The methodology was straightforward:
* Logged every inline suggestion Cline made in the IDE.
* Categorized the outcome: Accepted, Modified & Accepted, or Ignored/Rejected.
* Broke it down by language (TypeScript, Python, Go) and suggestion type (code completion, bug fix, refactor).
The initial aggregate number is... underwhelming. The overall acceptance rate sits at 34%. The breakdown is more telling:
* **Code completions** (e.g., finishing a line): ~65% acceptance. This is its strong suit, but it's also the lowest-value function.
* **"Bug fix" suggestions**: ~22% acceptance. Most were either irrelevant or would have broken something else.
* **Refactoring proposals**: ~15% acceptance. These were almost universally style changes we don't follow or overly complex "solutions" to simple code.
More concerning than the low rate is the noise floor. For every accepted suggestion, the team dismissed about four others, which is a tangible context-switching cost. I'm also tracking the latency of the suggestions; slower ones are almost always ignored as the developer has already moved on.
I'm curious if others have done similar tracking. Specifically:
* What acceptance rates are you seeing, and how are you measuring?
* Have you quantified the distraction cost of bad suggestions?
* Does the rate improve significantly after extensive prompt tuning, or is this just the ceiling?
I'll post the dashboard config and query logic if there's interest. The vendor's SLA promises "increased velocity," but I want to see the metric they'd use to prove it if we had it in a contract.
- Ray
- Ray
Quantifying acceptance rate is the right first step, but that 34% aggregate is a misleading average. You need to track the rate of *harmful* suggestions alongside it. A low-value code completion is fine, but a "bug fix" with a 22% acceptance rate implies 78% of those were noise or wrong. How many of those rejected fixes wasted developer time parsing them? That's a cost metric you're missing.
Your categorization is good, but the time dimension will likely show more. Did the acceptance rate for refactoring proposals improve as the tool learned your codebase patterns, or did it flatline? A static quarterly number buries that trend. You should also segment by individual developer if you can; variance there often points to differences in prompting style or areas of the codebase where the tool has less context.
I'd instrument the latency between a suggestion appearing and its final state - accepted, modified, or dismissed. If developers are spending more than a few seconds evaluating a suggestion only to reject it, that's active productivity drain, not just passive noise. That's a key Service Level Objective for any "co-pilot" that rarely gets measured.
metrics over vibes