After observing significant variance in the quality and utility of commit messages across my team's repositories, I initiated a controlled experiment six months ago. The objective was to standardize and elevate commit message quality by delegating their creation entirely to a coding assistant (specifically, GitHub Copilot Chat and Claude via CLI), with a strict protocol. The hypothesis was that this would not only improve the historical record but also enforce more disciplined, atomic commits by requiring the AI to generate a coherent message for each discrete change.
The core of the workflow is a pre-commit hook script that leverages `git diff --staged` to capture changes. This diff is then passed to the AI model with a carefully engineered prompt. The key was to move beyond simple summarization and demand a specific, structured format that aligns with conventional commit standards and provides actionable context for future investigation (e.g., during bisect). Below is the essential section of the script I developed and refined over the trial period.
```bash
#!/bin/bash
# .git/hooks/prepare-commit-msg
STAGED_DIFF=$(git diff --staged --no-ext-diff)
if [ -z "$STAGED_DIFF" ]; then
exit 0
fi
# Construct the prompt with clear instructions for structure and focus.
PROMPT="Generate a conventional Git commit message for the following staged diff.
Your output must be a single line for the subject, followed by a blank line, followed by a body.
Subject Line: Use the imperative mood, capitalize the first letter, and do not exceed 50 characters.
Body: Explain the 'why' and the context, not just the 'what'. Reference any fixed issues. Use bullet points if needed.
Focus on the change's intent and impact. If it's a refactor, state the motivation.
Here is the diff:
${STAGED_DIFF}"
# Using Claude via Claude CLI for this example. Similar setup works for Copilot.
AI_MESSAGE=$(claude "$PROMPT" --model=claude-3-sonnet-20240229)
# Overwrite the default commit message file with the AI-generated one.
echo "$AI_MESSAGE" > "$1"
```
The results were measured across several dimensions by analyzing the commit history from the six months prior to implementation and the six months after. A cohort analysis of repository contributors was also performed to control for individual variance.
**Quantitative Findings:**
* **Consistency:** 98% of AI-generated commits adhered to the conventional commit format, compared to a baseline of 45%.
* **Atomicity:** The average lines changed per commit decreased by approximately 35%, as the practice of staging logically grouped changes became necessary for the AI to produce a sensible message.
* **Bisect Utility:** In a simulated `git bisect` exercise on 10 known bugs, the clear AI-generated messages led to correct identification of the introducing commit 20% faster on average.
**Qualitative & Workflow Observations:**
* The requirement to stage a coherent diff for the AI acted as a forcing function for more thoughtful, incremental commits. The "commit early, commit often" adage was realized with consistent quality.
* The body of the messages reliably included rationale, which has proven invaluable during code reviews and refactoring efforts.
* A notable, unexpected benefit was the exposure of "diff hygiene." Poorly structured changes (e.g., mixing formatting edits with logic changes) resulted in confusing AI messages, which prompted immediate correction of the staging. This had a positive secondary effect on code quality.
* The primary cost is a slight increase in commit time (approximately 2-3 seconds for model inference). However, this is offset by the time saved not crafting messages manually and the reduced context-switching.
**Final Configuration & Checklist:**
After iterative refinement, my current protocol includes the following pre-commit checklist to ensure optimal input for the AI:
* [ ] Run linter/formatter to ensure diff is not polluted with style changes.
* [ ] Stage only the changes pertaining to a single logical unit of work.
* [ ] Verify `git diff --staged` is readable and focused.
* [ ] Review the AI-generated message for accuracy before finalizing the commit; it is an assistant, not an authority. Edits are made in roughly 15% of cases.
This systematic approach has transformed commit messages from an afterthought into a robust, searchable log of development intent. The consistency and depth of context have materially improved the efficiency of code archaeology and team onboarding.
Data > opinions
That's a fascinating approach, and I can see the immediate appeal. Using the AI to enforce the discipline of small, atomic commits is a clever side effect I hadn't considered before. It makes the commit message a forcing function, not just an afterthought.
My main hesitation would be around the training data's bias. The structured format you're prompting for is excellent, but I wonder if there's a tendency for the AI to overfit toward what it's seen in public repos, which often don't reflect the specific, internal business logic of a private project. Does it sometimes produce messages that sound correct but miss a subtle, domain-specific "why"? I've seen that happen with code generation, and I'd be curious if it translates to this use case.
—daniel
You've absolutely put your finger on the real risk here. That overfit to generic patterns is a huge issue, and it absolutely translates. I've seen it create a perfect-sounding "Fixed null reference in user validation" message that completely omitted the crucial detail that this only happened for legacy accounts imported during the Q3 migration - the exact "why" a developer would need six months later.
The saving grace, ironically, is that the forcing function of small commits makes the diff so tiny and focused that the AI has less room to hallucinate context. It's just looking at, say, three changed lines in one validation method. That helps, but it's not perfect. You still need a human to glance at the generated message and ask, "does this capture the actual business reason?" before hitting commit. It becomes a collaborative prompt-review cycle, not a full abdication.
Implementation is 80% process, 20% tool.