Skip to content
Notifications
Clear all

Walkthrough: Building a custom 'code smell' classifier with OpenClaw's base model.

1 Posts
1 Users
0 Reactions
2 Views
(@amandaj)
Reputable Member
Joined: 1 week ago
Posts: 148
Topic starter   [#19571]

In my ongoing evaluation of AI-assisted code review systems, I've found that while most commercial tools are proficient at catching generic bug patterns, they often lack the specificity required to enforce our team's particular architectural and code quality standards. A common "non-bug" but significant issue for us is the proliferation of certain custom patterns—what we internally term "legacy service patterns"—that we are actively trying to migrate away from. No out-of-the-box tool catches these contextual "code smells."

This led me to experiment with fine-tuning a base model to create a specialized classifier. After testing several platforms, I settled on using OpenClaw's publicly available `claw-7b-base` model for this walkthrough, due to its strong performance on code-related tasks and its suitability for instruction tuning. The goal was to produce a model that could scan a pull request diff and flag instances of our deprecated patterns with high precision, integrating directly into our CI pipeline.

The process involved several distinct phases:

1. **Dataset Curation & Annotation**
We extracted approximately 500 code snippets from our historical repository data, focusing on files touched in the last two years. Each snippet was labeled as either containing the "legacy pattern" or not. The key was to include not just the problematic line, but sufficient surrounding context (typically 5-10 lines before and after). The dataset was structured as a JSONL file for training.

```json
{"messages": [{"role": "system", "content": "You are a code review assistant. Analyze the provided code snippet and determine if it contains the legacy service pattern, characterized by direct instantiation of the 'LegacyServiceClient' and its 'call()' method."}, {"role": "user", "content": "```pythonndef fetch_user_data(user_id):n # ... some setup ...n client = LegacyServiceClient(n host=config.host,n port=config.portn )n response = client.call('get_user', user_id)n return parse(response)n```"}, {"role": "assistant", "content": "Yes, this code contains the legacy service pattern."}]}
{"messages": [{"role": "system", "content": "..."}, {"role": "user", "content": "```pythonndef fetch_user_data(user_id):n return NewServiceAdapter.get_user(user_id)n```"}, {"role": "assistant", "content": "No, this code does not contain the legacy service pattern."}]}
```

2. **Instruction Tuning with LoRA**
Given the relatively small dataset, full fine-tuning was impractical. I employed Low-Rank Adaptation (LoRA) to efficiently train the model. Using the `peft` and `transformers` libraries, the setup targeted the query and value attention matrices. The configuration emphasized a low rank (`r=8`) to prevent overfitting.

```python
from peft import LoraConfig, get_peft_model

lora_config = LoraConfig(
r=8,
lora_alpha=16,
target_modules=["q_proj", "v_proj"],
lora_dropout=0.05,
bias="none",
task_type="CAUSAL_LM"
)
model = get_peft_model(base_model, lora_config)
```

3. **Training & Evaluation Split**
We used an 80/20 train/test split. Crucially, the test set contained entirely *unseen* files and patterns not present in the training data to better simulate real-world performance. The primary metrics were precision and recall on the "positive" (pattern present) class, as false positives (incorrectly flagging clean code) would create unacceptable noise in our review queue.

4. **Inference Integration**
The final tuned model was wrapped in a lightweight FastAPI service. It accepts a diff string, breaks it into manageable hunks with context, and runs classification on each. The output is a structured JSON report suitable for our CI bot to comment on the PR.

**Initial Results & Observations**

After 3 training epochs, the classifier achieved the following performance on our held-out test set:
* **Precision:** 0.94
* **Recall:** 0.87
* **F1-Score:** 0.90

Manual review of the false positives revealed they were often in test files or involved refactored versions of the pattern that were, in fact, acceptable. This precision level is acceptable for our initial deployment. The recall miss (13% of instances not caught) primarily involved highly obfuscated or abstracted instances of the pattern, which we consider a acceptable trade-off for now.

The primary advantage of this approach over prompt engineering with a general-purpose model is cost and latency. The fine-tuned 7B model runs efficiently on a single GPU instance and provides consistent, rule-like behavior tailored to our codebase. The main cost was the initial, somewhat labor-intensive, dataset creation.

I am now planning a cohort-based analysis to measure its impact over the next quarter: tracking the reduction in new introductions of this pattern and surveying developer sentiment on the tool's feedback. The next step will be to expand the classifier to a second, related smell.

Has anyone else undertaken a similar project of training a highly specific, domain-aware code reviewer? I'm particularly interested in comparisons of dataset sizing strategies or alternative base models for this classification task.

— Amanda


Data > opinions


   
Quote