Skip to content
Notifications
Clear all

Has anyone quantified the time saved on writing documentation strings with Copilot?

1 Posts
1 Users
0 Reactions
23 Views
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
Topic starter   [#4144]

Everyone's raving about Copilot's ability to autocomplete entire functions, but the real, insidious productivity sink has always been writing the docstrings and comments that make that code comprehensible six months later. I've been skeptical of the "time saved" claims, especially for documentation, which requires actual intent and understanding.

So I ran a small, deeply unscientific experiment on a legacy Python service module I was refactoring. I wrote docstrings for 20 functions manually, timing myself. Then I reverted, let Copilot suggest them, and timed the editing required to make them correct and useful.

The raw suggestion was often dangerously generic. For a function parsing a specific, poorly-formatted log line, it would propose:

```python
def parse_log_line(raw_line):
"""
Parses a log line.

Args:
raw_line (str): The raw log line.

Returns:
dict: The parsed log line.
"""
```

This is worse than nothing—it's a lie. It implies a standard return type and doesn't capture the weird edge cases. But, it *does* give you the stub. The time wasn't saved in acceptance, but in not having to type the boilerplate `Args:` and `Returns:` sections. The win was in tab-completing the structure, not the content.

My rough numbers? Manual: ~90 seconds per function for thoughtful docs. With Copilot: ~45 seconds, but almost all of that was me fighting to correct its assumptions. The net saving was maybe 30%, but only because I was willing to aggressively rewrite its proposals. If you blindly accept, your documentation becomes a hollow artifact that actively misleads.

The real question isn't "time saved," but "time shifted." It moves the effort from typing to editing and fact-checking. In domains with highly conventional patterns (REST controller methods, CRUD operations), that shift might be beneficial. In the messy reality of legacy systems and edge-case logic, you're just trading one cognitive load for another. Has anyone else moved past the initial "wow" and tried to measure what it actually does for non-greenfield work?



   
Quote