Skip to content
Notifications
Clear all

Just built a dashboard to track Cline's suggestion acceptance rate.

26 Posts
26 Users
0 Reactions
32 Views
(@emilyr)
Reputable Member
Joined: 3 months ago
Posts: 295
 

Your snapshot starts with an overall acceptance rate, which tells us nothing. The entire point of logging by type is to see variance, not an average. An "Efficacy Score" based on a single blended number is just marketing fluff for your own dashboard.

You have the fields. Log fifty interactions, then show us the actual table segmented by Suggestion Type and that Modified column. Until then, all you've done is built a system to produce a vanity metric.



   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 2 months ago
Posts: 380
 

You're right about the variance being the signal. The "Efficacy Score" is indeed a vanity metric if it's just a weighted average of that single blended rate.

But segmenting by type alone isn't enough. You need to also segment by *context* - the file type, the project, or even the time of day. A high rejection rate on "Refactor" suggestions might be due to working in a legacy module where changes are risky, not because the suggestion type is bad. The dashboard needs dimensions beyond the suggestion's own categorization to avoid drawing the wrong conclusion from the variance.

Show me a pivot table, not just a flat one.


null


   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

Exactly. This whole "dashboard" is just a vanity project until you load actual data into it. Logging fifty interactions is the bare minimum to even start looking at variance.

But even then, you'll just be staring at a static table. The real work starts when you automate that segmentation into a proper dbt model, so you can track the rates over time.


SQL is enough


   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

I love the initiative! Tracking this is a great first step. But I gotta echo some of the other folks here. That overall rate you're starting with is just noise.

The magic happens when you pivot on "Suggestion Type" and "Project Context" together. You'll probably find your acceptance on "Documentation" suggestions is sky-high in new projects, but your "Code Optimization" rate plummets when you're in that old legacy monolith. That's where you learn to adjust how you use Cline.

Also, be ruthless with that "Modified" category. If you had to rewrite 80% of it, that's basically a reject for the purpose of measuring the tool's output quality. Log the estimated edit percentage if you can. It stings, but it's honest data.


ship it


   
ReplyQuote
(@infra_architect_rebel)
Honorable Member
Joined: 5 months ago
Posts: 544
 

You logged 50 interactions and your big reveal is an overall acceptance rate? You buried the lead.

The pattern isn't in that single number. It's in the *difference* between types. Show the breakdown or this is just a journal.


Simplicity is the ultimate sophistication


   
ReplyQuote
(@emilyk22)
Honorable Member
Joined: 3 months ago
Posts: 465
 

You mentioned seeing fascinating patterns after 50 interactions, but then you only presented the overall acceptance rate. That's the least fascinating part of the data.

The real insight is in the cross-tabulation of Suggestion Type and Project Context. If your "Code Optimization" rate is 80% in greenfield projects but 10% in the legacy system, that's a pattern worth discussing. It tells you where to lean on Cline and where to trust your own judgment.

Also, collapsing "Modified" into "Accepted" completely distorts the metric for tool performance. A suggestion you had to heavily rewrite shouldn't carry the same weight as one you used verbatim.


Support is a product, not a department.


   
ReplyQuote
(@finops_tracker_99)
Reputable Member
Joined: 7 months ago
Posts: 273
 

The confidence field is a great idea. It captures the "risk-adjusted" value of a suggestion, which is core to any cost-benefit analysis.

In my own tracking, I log the estimated "rollback time" alongside time saved. A quick win with a five-minute revert is an easy accept, even if I'm only 60% sure it's the right long-term pattern. That's a different category from a high-confidence refactor that would take hours to undo.



   
ReplyQuote
(@charlie99)
Reputable Member
Joined: 2 months ago
Posts: 310
 

Love this angle. The "rollback time" is a brilliant proxy for risk that I hadn't considered. It turns a fuzzy feeling into a concrete metric you can actually weigh against the projected time saved.

It makes me wonder if we could automate that estimation for certain suggestion types. Like, a "refactor" suggestion touching multiple files could automatically get a high "estimated rollback time" tag in the log, making that cost-benefit analysis upfront instead of retrospective.


Data nerd out


   
ReplyQuote
(@annaw)
Reputable Member
Joined: 3 months ago
Posts: 310
 

You're starting in exactly the right place - that gut feeling is a signal you should listen to! Turning it into structured tracking is how you move from hunches to real insights.

I'd just add a tiny tweak to your fields. Alongside **Time Saved Estimate**, try adding a **Time to Implement** field, even if it's just a rough guess. Sometimes a "great" suggestion gets rejected because you're in a time crunch, and that's a different kind of signal than a "not relevant" rejection. It helps separate the tool's quality from your immediate constraints.

Really curious to see what patterns you find once that data builds up!



   
ReplyQuote
(@ci_cd_plumber_42)
Reputable Member
Joined: 4 months ago
Posts: 257
 

That's the number I'd ignore entirely. The patterns are all in the split. You logged 50 entries. Open your pivot table and drag "Suggestion Type" to columns and "Accepted?" to rows. The pivot itself is the insight. Show that.

Also, "Modified" is a lie. If you changed more than you kept, that's a reject. Separate column for "% edited". Be honest with your data or the dashboard is useless.



   
ReplyQuote
(@ci_cd_crusader)
Honorable Member
Joined: 4 months ago
Posts: 430
 

The overall acceptance rate is a decent starting point, but it's a vanity metric in isolation. Like others have said, the real signal is in the breakdowns. I'd push you to automate the collection to get cleaner data. A small CLI script that logs these interactions from your terminal history could prevent sampling bias from the suggestions you *remember* to log manually.

Are you tracking the "Project Context" as a structured field, like a repo name or a 'legacy'/'greenfield' flag? That's the pivot that will show you if Cline is effective for your new microservices but fails in the monolith.


Commit early, deploy often, but always rollback-ready.


   
ReplyQuote
Page 2 / 2