Skip to content
Notifications
Clear all

First-time evaluator: what metrics should I track in a trial?

2 Posts
2 Users
0 Reactions
25 Views
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
Topic starter   [#19680]

I'm in the early stages of evaluating Semgrep for our data platform team. While I'm comfortable benchmarking SQL execution times or dashboard load speeds, static analysis tool evaluation is new territory. I want to move beyond "did it find a bug?" to more systematic, operational metrics.

For those who have run a structured trial, what did you measure? I'm thinking along two axes:

**Effectiveness & Accuracy**
* **Precision/Recall:** Of the rules in a standard security rule pack (e.g., `p/ci`), what percentage of findings were true positives vs. false positives? Manually auditing a sample is necessary here.
* **Critical Findings:** Raw count of high-severity, exploitable issues discovered that were previously unknown.
* **Rule Customization Success Rate:** For custom rules written for our internal frameworks (like specific DBT macro patterns), what was the false positive rate?

**Operational & Performance**
* **Scan Duration:** Baseline scan time for our codebase, tracked against incremental scan times. This impacts CI/CD integration feasibility.
* **Integration Effort:** Time to get a baseline scan running locally and in CI (GitHub Actions/GitLab CI). Configuration complexity is a factor.
* **Learning Curve:** Time for a data engineer (proficient in SQL/Python but new to Semgrep) to write a simple, effective custom rule.

My initial setup for tracking scan time looks like this:
```bash
time semgrep scan --config auto --metrics=off -q > findings.json
```

I'm particularly interested in how you quantified the "softer" aspects, like reduction in code review burden or the efficiency of the rule syntax for team adoption. Are there any standard benchmarks or community-agreed-upon KPIs for tools like this?



   
Quote
 danw
(@danw)
Reputable Member
Joined: 3 months ago
Posts: 387
 

Skip precision/recall on the standard packs. Vendor packs are a solved problem - they'll perform as advertised. Your time is better spent on custom rule efficacy.

Track the "time to fix" loop. How long from a developer receiving a finding to a merged commit that resolves it? That's your real operational metric. If it's high, your integration is broken.

Also measure noise in CI. Does it break builds with low-severity stuff? If yes, teams will start disabling it.



   
ReplyQuote