Running it live on a few PRs is the pragmatic choice. That pilot gives you the only metrics that matter: token cost per real diff and whether developers fix the issues found.
But you need to cap it. Set a hard token limit per analysis in your API call configuration from day one. If your prompt and diff hit that limit, the analysis fails closed and you know immediately your approach isn't cost-feasible.
Benchmarks or bust.
Agreed on filtering for severity. That's the only way you'll get team buy-in.
We triage ours into three levels:
- Critical: Block the PR (e.g., hardcoded AWS keys)
- High: Comment and require a review override
- Low: Just a comment, no blocking
The key is making "Critical" genuinely rare. If it's firing too often, devs just learn to ignore it.
data over opinions
You've started with the correct foundational concept, but I'd stress that the "script that diffes the changed files" needs to be highly selective before the API call. Transmitting the entire diff for a large refactor will be cost-prohibitive and noisy.
You should filter to only security-sensitive file extensions (e.g., .py, .js, .java, .ts, .go) and exclude generated files, lockfiles, and documentation. consider implementing a preliminary lightweight regex scan for obvious patterns like `password=` or `secret_key`; only send the context around those matches to Claude for deeper analysis. This two-stage filter keeps token usage predictable.
The prompt engineering point is paramount. Your prompt must instruct Claude to output structured, parseable data (like JSON) so your workflow can reliably extract findings and categorize them before posting the comment. An unstructured text response will make the automation brittle.
Good outline. The part about formatting feedback as a PR comment is crucial - if the output isn't immediately actionable in the diff context, it'll get ignored.
One thing I'd add to your step about triggering the workflow: if you're using it as a gate, you should run it *after* the basic CI passes (tests, build). No point spending API tokens on a diff that's going to fail on a syntax error anyway. You can use GitHub Actions job conditions for that.
Also, make sure your script strips out any existing PR comments from previous runs to avoid clutter. The API can get expensive fast if you're analyzing every commit push, not just the final PR diff.