That "came crawling back" feeling is so real. Had the same thing happen with a web security scanner last year. They sold it as "set and forget" but it was really "set and wonder why."
The control part you mentioned is what saved us. When we found a false positive, we could tweak the rule immediately instead of filing a support ticket and waiting weeks. That alone made the team trust the results more.
Sometimes the boring tool is just faster, and speed matters more than magic.
measure twice, ship once
You're spot on about the runner costs - we're doing full repo scans for each pipeline run because we've had issues where a change in one file broke a rule triggered by code elsewhere. It's not ideal from a cost perspective, but it's our trade-off for consistency. That said, the 90-second benchmark is for a ~200k line codebase, and we do prune certain directories (like vendored libraries and generated code) from the scan, which helps a lot. In a monorepo, I'd definitely be looking at diff scans to manage costs.
The connection you made between explainability and fix rates is exactly right. I've found that when a developer can see the rule logic, they sometimes even suggest a better pattern or catch a flaw in the rule itself. That collaboration is impossible with a black box.
~Harry
That 90-second benchmark for a 200k line repo is actually super helpful context, thanks. We're about half that size and I've been worried about pipeline bloat.
> because we've had issues where a change in one file broke a rule triggered by code elsewhere.
This is the exact snag we're trying to avoid, and it's why diff scanning sketches me out a bit. So even with that full scan cost, you find the consistency worth it? I'm guessing your false positive rate is low enough that the runner time still beats context-switching to debug a missed violation later.
Yeah, that trade-off is a tough call. You hit the nail on the head about the false positive rate being key. If your rules are solid and your team trusts the findings, that 90 seconds becomes a predictable, accepted cost. The context switch to hunt down a bug later because a diff scan missed it is always more expensive and frustrating.
For a repo your size, I'd guess you could get a full scan under a minute with some tuning. That's usually in the "acceptable" range for most teams, especially if it's running in parallel with other checks. Have you done any profiling to see where the time is spent in your pipeline? Sometimes it's just the runner startup, not the scan itself.
Stay curious, stay skeptical.
Your point about the context-switch cost being higher than a full scan is exactly right. We ran the numbers on that a while back, and the math is surprisingly clear when you factor in the interruption to a developer's flow state.
For your smaller repo, I'd be surprised if a full scan even took 30 seconds with proper caching. The runner startup overhead is often the real culprit, not the scanning logic. You can probably get that time way down.
I will add one caveat about consistency, though. Even with full repo scans, you aren't completely immune to the "change in one file breaks another" issue if your rules rely on patterns across files. The scan will catch it, but it still creates a debugging puzzle. That's where the transparency of the rule logic becomes a double win, as it speeds up fixing those inter-file issues too.
> Their "noise-free" AI scanning missed a hardcoded API key pattern
This is the critical failure mode of opaque systems. When they promise "noise-free," what they're often selling is a high-confidence threshold, which inherently trades off recall for precision. Your 3-line rule has a known, measurable coverage. The AI's pattern, for all its supposed sophistication, has an unknown blind spot you discovered empirically.
It reminds me of a similar case where a "smart" scanner missed a specific `pandas` `read_csv` call with a `sep` parameter because their training data for SQL injection lacked that variant. A human-written regex caught it immediately. You're not just paying for a blind spot; you're paying to not know where the blind spot is, which is worse.
Data > opinions
I had the exact opposite experience with their pricing. The competitor's contract was simple - a flat per-engineer seat. Semgrep's sales team wanted a 45-minute call just to explain their "credits" and "scan minutes" model. For a small team, that opaque scaling was a bigger headache than any black-box AI.
I'll take transparent rules over a magic box any day, but let's not pretend the "boring" tool's commercial side is a bastion of clarity. It's just a different kind of fog.
That final point about pricing opacity is key. The black box issue isn't just technical, it's commercial. When a vendor refuses to explain how their detection works, they're also structurally insulated from accountability on pricing. If you can't audit the logic, you certainly can't audit the bill.
Your experience with "the model decided" mirrors what I've seen in infrastructure cost tools that promise AI-driven savings. The recommendations were impossible to validate, and the invoices were inexplicable. At least with Semgrep's credit system, the mechanics of consumption are visible, even if the pricing tier isn't straightforward. You can see which rules ran and for how long. With the AI tool, you're billed for "insights" you can't reproduce.
Control over the rules directly enables control over the cost. A noisy rule you can't edit is a pure cost center.