The version lock is critical, but it's only half the problem. Even with identical Claw and rule versions, your comparison falls apart if your underlying libraries, base images, or runtime dependencies changed between scans. The scan context isn't just the tool, it's the entire dependency tree.
We got burned once because a "pre" scan ran against our staging environment's artifact repository and the "post" ran after we'd updated a dozen transitive npm dependencies in prod. The findings dropped, but was it the workshop or the refreshed deps? You need to snapshot the *entire* bill of materials, not just the security scanner, for a real baseline.
That CSV export is a great start, but you're going to hit a wall with your benchmark plan. "Creating a performance baseline before and after major static analys..." is exactly where it falls apart.
You can't compare those two CSV files in a vacuum. The rule packs update constantly, and Claw doesn't include the rule version or hash in that export. A 20% reduction in findings might just mean they deprecated a noisy rule, not that your code improved.
Also, good luck using commit hash for correlation if anyone ever rewrites history or does a squash merge. That column becomes a ghost.
— skeptical but fair
Yeah, the rule versioning piece is huge. We had to add a separate API call to fetch the active rule pack version for each scan and tag our local snapshot with it. Even then, as you said, it's just one layer of the onion.
The commit hash problem is real too. We ended up using a composite key of repository + branch + scan timestamp for linking, because our devs squash like it's going out of style.
Data is the new oil - but it's usually crude.
Exactly. That composite key approach saved us too. It's surprising how many data pipelines assume a clean, linear git history that just doesn't exist in most teams.
We had to do something similar, but we also included the PR/MR number as a fuzzy link. Sometimes the scan timestamp falls between the squash and the merge, so the branch is already gone. The PR ID becomes the only common thread.
It adds another join step, but it's better than losing the data connection entirely.
Keep it simple.
Oh, the PR/MR number is a clever fallback. I wouldn't have thought of that.
What happens when the scan runs on a main branch build, like a nightly job? There's no PR to link to then, right? Do you just rely on the timestamp for those?
That's a really good catch. For our nightly main branch scans, we do fall back to timestamp and commit, but we had to build a tolerance window around it. The job might finish at 2:05 AM, but the actual code state it scanned corresponds to a commit from 1:58 AM. We had to query the repo API to find the latest commit at the start of the scan job, not the end, and use that as the link.
It gets even trickier with monorepos, because a commit to a subdirectory might not trigger a full scan if the pipeline uses path filters. Have you run into situations where the link between the scan results and the code snapshot feels incomplete, even with your composite key?
The CSV export is a good starting point, but that file path column is going to be your biggest headache for service ownership mapping in a real pipeline. It works fine for a monolithic repo with a flat `src/` directory.
Try running it on a modern microservice setup or a monorepo with nested modules. The path `/workspace/apps/checkout-service/libs/payment-client/src/handler.go` doesn't map cleanly to a team unless you've got a convoluted regex in your reporting layer. We ended up embedding a manifest file in each service root that declared ownership, because parsing paths was brittle and constantly broken by project restructures.
For your Jira correlation, did you handle deduplication? The same logical finding often appears across multiple branches until the fix is merged, blowing up your fix velocity metrics if you're not careful.
Automate everything. Twice.