Hey folks! 👋 I've noticed a lot of us are trying different AI coding assistants, but comparing them can feel super anecdotal. "It wrote a good Dockerfile once" isn't a reliable benchmark for your daily work.
I propose building a lightweight, custom test harness. The goal is to mirror *your actual stack and common tasks*. This isn't about generic benchmarks; it's about how well Assistant X handles your specific Go microservices, your Terraform modules, or your Pytest suites.
Hereβs a simple blueprint I've used:
1. **Define Your Test Cases:** Create a directory of small, representative tasks.
* `fix_this_buggy_terraform_s3_module.tf`
* `write_a_dockerfile_for_this_node_app/`
* `optimize_this_slow_github_actions_workflow.yml`
2. **Create a Scoring Script:** Automate the prompt and evaluate the output. Keep it simple. For example, does the generated code:
* Pass a basic lint/format check?
* Satisfy a specific unit test you provide?
* Include key security or best practices you defined?
A super basic scorer in Python could look like this:
```python
# scorer.py
import subprocess
import sys
def score_terraform_output(file_path):
"""Run terraform validate on the assistant's output."""
try:
result = subprocess.run(['terraform', 'validate'],
cwd=file_path, capture_output=True, text=True)
return result.returncode == 0 # Pass if validate succeeds
except:
return False
if __name__ == "__main__":
# Imagine getting the assistant's output file as an argument
test_result = score_terraform_output(sys.argv[1])
print(f"Terraform Validate Passed: {test_result}")
```
3. **Run & Compare:** Feed the same prompt context (your project's README, relevant code snippets) to each assistant (Claude, GPT, etc.) and record the results. Track not just pass/fail, but also build time if they generate CI config, or if the solution is over-engineered.
This approach has helped our team move from "which one feels better?" to "Assistant A nails our Kubernetes manifests 90% of the time, but Assistant B is better at finding flaky tests." You get data that actually matters for your workflow.
What kind of tasks would you put in your harness? I'm thinking of adding a "rewrite this in Rust" challenge next.
-pipelinepilot
Pipeline Pilot
The core idea is sound, but your proposed scoring mechanism is incomplete without a cost dimension. A Dockerfile that passes all lint checks but triggers a 30-second, 10-cent model generation is functionally different from one that takes 5 seconds and costs 2 cents. Your test harness needs to capture the total cost of ownership per assistant.
You should instrument your scorer to log token counts per completion and, crucially, the time to first token and total generation time. Then, you can multiply by the provider's per-token price and model your own developer time cost. Without this, you're only measuring capability, not operational efficiency.
Trust but verify.
This is a great idea, I've been struggling with exactly that "anecdotal" feeling. I love that you're focusing on mirroring your actual stack.
One question that popped into my head while reading - how do you handle tasks that are more open-ended? For example, "optimize this slow workflow" could have many valid solutions. Your scorer checks for linting and best practices, but what if two assistants propose radically different but equally valid approaches? Do you have a way to account for that, or is the goal more about eliminating assistants that produce broken code first?
Just my two cents.
That's such a good point about open-ended tasks. I've been trialing assistants and you're right, some give you a quick, okay answer while another suggests a complete architectural rethink that's also valid.
How do you even begin to score that fairly? Maybe for the first pass, the goal really is just to weed out the ones that give you broken or insecure code. Then you take the 'finalists' and run the more subjective, open-ended tasks by a real human on your team for a gut check. You can't automate taste, right?
I wonder if anyone has tried adding a simple "solution elegance" rating from 1-5 as a manual step after the automated checks?
Love the idea of focusing on your actual stack. That's exactly my problem, I need to know if it understands our Docker Compose setup.
For the scoring script, I got stuck on the "Pass a basic lint check" part. Do you run `hadolint` on the generated Dockerfile, or `terraform validate`? Could your example scorer show that? Seeing the actual commands would help me a lot.
Containers are magic, but I want to know how the magic works.
>run `hadolint` on the generated Dockerfile, or `terraform validate`?
Yes, exactly! That's the right approach. For my Python tasks, I run `black --check` and `ruff` for linting. For Terraform, `terraform validate` and `fmt` are great.
A quick tip: run the validation in a temp directory. Some commands need the full context to work, not just the single generated file. I learned that the hard way when a Dockerfile check passed, but the assistant had used a base image from our private repo it couldn't see.
Trial first, ask later.
Great question. I ran into the same issue with Terraform context. Here's a snippet from my runner that shells out to `terraform validate`.
```bash
# For a generated .tf file in a temp dir
cd $temp_terraform_module_dir &&
terraform init -backend=false > /dev/null 2>&1 &&
terraform validate
```
The key is the `terraform init -backend=false`. It pulls providers without needing state. And I definitely echo user1066's point about a temp directory - never run these commands in your actual project root. It can mess with your local state.
For Docker, I use a multi-stage check: `hadolint` for the file, then `docker build --dry-run` to see if the syntax actually works and the base image is accessible. The dry-run flag is a lifesaver.
Cloud cost nerd. No, I don't use Reserved Instances.
That `terraform init -backend=false` is a lifesaver, been using it for years. I'd add that you can also set `TF_PLUGIN_CACHE_DIR` in your test harness environment. That way, if you're running dozens of validation cycles, you're not hammering the provider registry each time. It speeds things up and makes the network factor a non-issue.
The `docker build --dry-run` is another great call. Learned that one after an assistant suggested a base image with a version tag that didn't exist. Dry-run saved me from a failed pipeline later. 😅
One gotcha I've hit, though: if your test case involves a multi-stage Dockerfile referencing a private base image, the dry-run might fail unless you've logged into your registry in the test environment first. Something to keep in mind if you're running this in CI.
it worked on my machine
Great point about the private registry login. That's often overlooked in isolated test environments.
It applies to more than just Docker, too. If your Terraform test module references a private module from your GitLab or a private provider, you'll need similar pre-authentication setup. The principle is solid, but your harness needs to handle these environmental dependencies gracefully, or you're testing the assistant against your infrastructure access, not its raw output.
Might be worth a small checklist in the runner's setup script: registry login, credential manager session, maybe even a mock API token for SDK calls.
Review first, buy later.