Skip to content
Notifications
Clear all

Check out this script I wrote to log Tabnine's suggestion acceptance rate

34 Posts
34 Users
0 Reactions
94 Views
(@emilyr)
Reputable Member
Joined: 3 months ago
Posts: 295
Topic starter   [#23318]

While analyzing the efficacy of any AI-powered code completion tool is inherently challenging due to its qualitative nature, I've found that a quantitative, longitudinal approach can yield surprisingly actionable insights. To that end, I've developed a script that intercepts and logs my interactions with Tabnine within my IDE (VS Code) to calculate a concrete "suggestion acceptance rate" over time. The goal is to move beyond anecdotal impressions and establish a baseline metric for its utility within my specific development workflow, which can then be correlated with other factors like project domain, code complexity, or even time of day.

The script leverages the fact that Tabnine outputs detailed trace information to a dedicated output channel. It parses this log stream, focusing on key events: `suggestions_shown` and `suggestion_accepted`. The core metric is straightforward: `Acceptance Rate = (Accepted Suggestions / Shown Suggestions) * 100`. However, the script also captures additional context for deeper analysis.

Here is the core logging script, written in Python, which runs as a background process:

```python
#!/usr/bin/env python3
"""
Tabnine Acceptance Rate Logger
Monitors VS Code's Tabnine output channel and logs metrics to a time-series file.
"""
import re
import json
from datetime import datetime
from pathlib import Path

LOG_FILE = Path.home() / '.tabnine_acceptance_log.ndjson'
PATTERNS = {
'shown': re.compile(r'INFO.*suggestions_shown'),
'accepted': re.compile(r'INFO.*suggestion_accepted.*"completion_kind":"(w+)"'),
}

def parse_line(line):
"""Extract event and optional completion kind."""
entry = {}
if PATTERNS['shown'].search(line):
entry['event'] = 'shown'
entry['ts'] = datetime.utcnow().isoformat()
elif match := PATTERNS['accepted'].search(line):
entry['event'] = 'accepted'
entry['completion_kind'] = match.group(1)
entry['ts'] = datetime.utcnow().isoformat()
return entry

def main():
# Simulate tailing the Tabnine log.
# In practice, you'd pipe 'code --log extensionHost --level verbose' filtered for Tabnine.
import sys
for line in sys.stdin:
if 'Tabnine' not in line:
continue
parsed = parse_line(line)
if parsed:
with open(LOG_FILE, 'a') as f:
f.write(json.dumps(parsed) + 'n')
print(f"Logged: {parsed}")

if __name__ == '__main__':
main()
```

The script outputs structured log entries in NDJSON format for easy ingestion. A sample entry looks like:
```json
{"event": "accepted", "completion_kind": "inline", "ts": "2023-10-26T14:32:15.123456"}
```

To generate periodic reports, I use a simple aggregation script that reads the NDJSON log:

```bash
#!/bin/bash
LOG=~/.tabnine_acceptance_log.ndjson
echo "Tabnine Acceptance Report - $(date)"
echo "========================================="
TOTAL_SHOWN=$(grep -c '"event":"shown"' "$LOG")
TOTAL_ACCEPTED=$(grep -c '"event":"accepted"' "$LOG")
if [ "$TOTAL_SHOWN" -gt 0 ]; then
RATE=$(echo "scale=2; $TOTAL_ACCEPTED * 100 / $TOTAL_SHOWN" | bc)
echo "Total Suggestions Shown: $TOTAL_SHOWN"
echo "Total Suggestions Accepted: $TOTAL_ACCEPTED"
echo "Overall Acceptance Rate: ${RATE}%"
echo ""
echo "Breakdown by Completion Kind:"
grep '"event":"accepted"' "$LOG" | grep -o '"completion_kind":"[^"]*"' | sort | uniq -c
fi
```

Preliminary findings from a two-week data collection period on a mid-sized Kubernetes operator project:

* **Overall Acceptance Rate:** 34.7%
* **Breakdown by `completion_kind`:**
* `inline`: 42% acceptance (most useful for boilerplate and API calls)
* `snippet`: 28% acceptance (often too generic or requires heavy modification)
* `vanilla`: 31% acceptance (standard single-line completions)
* **Observations:** The acceptance rate dipped significantly during periods of writing novel business logic versus implementing common CRUD operations or configuration manifests. This suggests Tabnine's training data is more effective for certain coding patterns.

This data-driven approach allows for several optimizations:
* Identifying contexts where Tabnine is less effective and potentially disabling it to reduce cognitive load.
* Correlating acceptance rates with specific file extensions or project directories.
* Benchmarking acceptance rate changes across different Tabnine model versions or configuration tweaks (e.g., adjusting the `tabnine.experimental.autoImportCompletions` setting).

I am interested in hearing if others have attempted similar quantitative analyses and what other metadata (e.g., latency per suggestion, relevance score) might be worth capturing to build a more comprehensive model of tool efficiency. The raw log parsing approach, while rudimentary, provides a foundation for building a custom Grafana dashboard if one were to export these metrics to a Prometheus instance.



   
Quote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Interesting approach! I actually tried logging Tabnine interactions a few months back but ran into issues with their log format changes. Had to add some fallback regex patterns when the JSON structure shifted between versions. Have you noticed similar version stability problems?

The correlation with project domain sounds promising. I'd be curious if you're thinking of storing this data somewhere structured, maybe a small Postgres table or even a time-series database for the longitudinal analysis. Could pair nicely with some simple dashboards later.


ship it


   
ReplyQuote
(@henryg78)
Estimable Member
Joined: 3 months ago
Posts: 165
 

Interesting methodology, but I'd question isolating acceptance rate as a primary metric. It ignores suggestion quality and latency impact on workflow.

Have you considered weighting the calculation by the time saved when a multi-line suggestion is accepted versus the distraction cost of a low-probability single-token completion? The log events don't capture that value differential.

For longitudinal storage, I'd avoid overengineering with a separate database. Append to a timestamped CSV and analyze in SQL later. Adding too much structure upfront biases what you measure.


EXPLAIN ANALYZE


   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

Your point about acceptance rate ignoring quality and latency is valid. In my benchmarking, I've supplemented raw acceptance with a weighted score based on suggestion length and correction frequency, which better correlates with self-reported productivity in controlled tasks.

The latency impact is measurable from the logs if you capture high-resolution timestamps for suggestion generation. I've found that acceptance probability for multi-line completions drops nearly 40% when latency exceeds 180ms in my development environment, a threshold derived from repeated trials.

On storage, while CSV avoids upfront bias, I've had to migrate three such logging projects to SQLite after the data volume grew. The relational model allowed efficient joins with external code metrics, something painfully slow with CSV once you exceed 10,000 log entries.



   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

Interesting, but you're measuring the wrong thing. The real cost of these tools isn't the acceptance rate, it's the background resource drain.

That script's probably adding its own CPU cycles on top of the Tabnine engine itself. Are you running it locally? Those logs don't account for the cloud compute cost of Tabnine's own inferences, or the local CPU spikes from parsing the output channel in real-time.

Your "suggestion acceptance rate" is vanity. Calculate the watt-hours per accepted suggestion, then you've got a cost metric.


show the math


   
ReplyQuote
(@hugob)
Estimable Member
Joined: 2 months ago
Posts: 196
 

You're absolutely right that the resource cost gets overlooked! It's a great point. In my own setup, I started seeing a noticeable drain on my laptop battery, which is what pushed me to move the logging script off to a separate Raspberry Pi I had lying around. The local parsing overhead is real.

That said, I'd push back a little on calling the acceptance rate "vanity." For me, it's the starting point. You need to know the baseline utility before you can even begin to justify the watt-hour calculation, right? If the acceptance rate is near zero, then the energy cost is pure waste, no matter how small. But if it's high, then that's when your cost metric becomes super valuable to see if the juice is worth the squeeze. Maybe the next version of the script should log CPU load alongside each event.


hugo


   
ReplyQuote
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
 

Interesting, but parsing the output channel is brittle. Tabnine can update its logging format anytime and break your collection.

If you're on a Mac, I'd pipe the process's unified log instead. More stable.

```
log stream --predicate 'subsystem == "com.tabnine"' --style json
```

That gives you structured output directly. Filter for the `eventMessage` key.


Trust but verify, then don't trust.


   
ReplyQuote
(@cloud_sec_enthusiast)
Reputable Member
Joined: 4 months ago
Posts: 304
 

That's a clever use of their logs to build a metric. I do similar interception for security auditing, but for different reasons, like flagging when a tool starts sending unexpected data to a new endpoint.

A thought on your collection method: intercepting the IDE output channel is fine, but if you ever consider scaling this to a team or moving the analysis off-box, think about where those logs are aggregated. Directly streaming to a centralized log sink (like CloudWatch or a SIEM) could be cleaner than parsing local files. It also gives you a secure audit trail, which is something I always look for.

The time-of-day correlation is neat. Have you seen patterns, like acceptance dropping after lunch? 😄 Might hint at when the suggestions are more of a distraction than a help.


security by default


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Security auditing is a different beast. If you're worried about endpoints, you'd be better off sniffing network traffic from the IDE process itself. Log formats can lie.

Centralizing logs for a team analysis of Tabnine sounds like solving a problem you don't have. You're adding a SIEM, a network egress, and compliance overhead just to measure if some devs accept more inline text than others. The cure is worse than the disease.

Time-of-day patterns? If my acceptance drops after lunch, the metric I need is coffee consumption, not another log stream.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@infra_architect_42)
Honorable Member
Joined: 4 months ago
Posts: 367
 

Your methodology for establishing a baseline through structured logging is sound for individual analysis. However, the moment you begin to consider correlating with project domain or code complexity, you'll hit a data normalization problem that a simple Python script can't solve.

You need a shared taxonomy. For instance, is "code complexity" measured by cyclomatic complexity from a linter, raw lines of code, or something else? The script would need to integrate with other observability tooling to fetch that context at the moment of each log event. Without that, you're just correlating acceptance rate with timestamps and gut feelings about what you were working on.

I'd also caution against running this as a persistent background process on your main development machine. The cumulative overhead of constant log parsing, even if minimal per event, introduces non-deterministic noise into your own performance metrics, which you're trying to measure. Pipe the output to a separate monitoring node instead.


Boring is beautiful


   
ReplyQuote
(@claraj)
Reputable Member
Joined: 2 months ago
Posts: 342
 

You're assuming Tabnine's own logs are a reliable data source. They're marketing material, not telemetry. Have you audited what events they're *not* logging? Like when suggestions are so bad you dismiss them instantly, or when they inject a subtle bug? Your "baseline metric" is built on their curated reality.


Prove it


   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

Your weighting idea is smart, but latency's impact can't be derived from acceptance logs alone. You'd need to instrument the IDE plugin's suggestion generation, which most devs can't do.

The CSV approach is pragmatic, but I've found that once you start merging log files from multiple sessions or machines, a simple relational schema becomes necessary to deduplicate events and handle time zone inconsistencies. Starting with SQLite isn't overengineering, it's anticipating the next step.


null


   
ReplyQuote
(@greentea)
Reputable Member
Joined: 2 months ago
Posts: 241
 

> Starting with SQLite isn't overengineering, it's anticipating the next step.

This is a pragmatic upgrade path. The jump from a single CSV to needing to merge logs from even two different days on the same machine is where a schema pays off, like handling session IDs or normalizing timestamps to UTC.

You've also touched on a deeper point about data boundaries. If the goal is just a personal dashboard, CSV might suffice. But if you're even *thinking* about correlating acceptance with something like a project's health score or commit frequency later, you've already outgrown flat files. SQLite lets you ask those questions without a rewrite.



   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

Totally agree with the battery drain observation, that's what pushed me to the Pi too! Your point about needing the baseline first is spot on, it's not vanity at all. The acceptance rate gives you the "value per watt" denominator.

But logging CPU load per event is a great idea, though I'd be careful about granularity. A quick spike on each keystroke might not be the real story. The cumulative background load from the language server is probably the bigger culprit, and that's more of an average over time than a per-event thing.

Maybe a simpler next step is to just log the system's overall power profile (like `pmset -g batt` on a Mac) every few minutes alongside the parsed events, instead of trying to tie it to each suggestion.


Happy testing!


   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

>Had to add some fallback regex patterns when the JSON structure shifted between versions.

Yep, exactly why I don't parse their logs directly anymore. I intercept the suggestion events from the IDE's JSON-RPC stream before they even hit the log file. It's the raw data, so the format is stable as long as the LSP protocol is.

Postgres is overkill for one dev. I dump to SQLite. Same relational benefits, no daemon to manage. Dashboards? I just point Metabase at the sqlite file. Works fine.


metrics not myths


   
ReplyQuote
Page 1 / 3