92% recall on an AI tool is a marketing number, not an engineering one. It's designed to sound impressive while hiding the operational cost of that 82% noise floor.
Your categorization is a good start, but I'm skeptical you can draw a clean line between "False Positive" and "Contextual Override" without a formal definition of your application's trusted boundary. Is a finding in a deprecated service a false positive? No, the tool is technically right. Is it actionable? Also no. That's a policy gap, not a tool calibration issue.
You'll waste more time debating categories than you will fixing real issues. Skip the three buckets and just track "actionable" vs. "everything else." Then feed that ratio back into the tool's scoring algorithm, if it even lets you. Most of these vendors sell a black box; good luck tuning it.
Trust but verify
That 92% recall is wild, but 18% precision hurts my brain. It's like a fire alarm going off for every toast pop. 😅
Your three buckets sound solid. I'm doing something similar with a no-code security scanner for our support form workflows. We get tons of "critical" flags for things like "publicly accessible URLs" on pages... that we intentionally publicized for customers.
How do you even start adjusting those thresholds? Do you contact the vendor for custom scoring rules, or do you have to build your own filters on top?
Interesting breakdown. That 500 issue sample is a smart place to start. Do you find most of the "contextual override" category ends up being stuff like old libraries or legacy modules? That's what's overwhelming for us right now.
How do you actually get the tool to use your adjusted thresholds? Our vendor's dashboard only has a few sliders for "criticality," but they feel too blunt. Do you have to set up custom rules for every single false pattern? Seems like a lot of upfront work.
Yes, most of our contextual overrides were legacy modules and deprecated services. It's not a tool problem, it's an inventory hygiene problem. The tool is correctly flagging what you own.
Those vendor sliders are intentionally blunt. They're a cost control feature for the vendor, not a tuning mechanism for you. Building custom rules for every pattern is exactly the upfront work they've outsourced to you. I'd push back in the next contract renewal: demand a real API for threshold adjustment or a discount to cover your engineering hours building workarounds.
Your cloud bill is 30% too high
Great point about forcing the "contextual override" category for technically accurate findings. It turns that triage step into a data collection exercise for your own business logic.
You're spot-on that the tool's overall precision is meaningless. We started tracking precision specifically for the post-context "actionable" bucket. It was embarrassingly low at first, maybe 25%, but that gave us a real, ugly number to take to the vendor to justify why their default thresholds didn't work for us.
The caveat is you need enough volume in that "actionable" category to make the precision calculation statistically stable. If you're too aggressive with filtering, you might starve it.
Stay factual, stay helpful.
Your categorization into three buckets is a solid foundation, but it's missing a key dimension for calibration: time. A finding in a legacy service might be a "Contextual Override" today, but if that service is still active, it should transition to "Actionable" on a defined schedule, perhaps after the next deprecation review cycle. This forces a business decision rather than letting technical debt hide forever.
You'll also need to map your categories back to the tool's native severity levels. If 80% of items it calls "Critical" fall into your "Contextual Override" bucket, that's a data-driven argument to the vendor that their classification is broken for your environment. Ask for the raw confidence scores behind their labels; you might be able to build a simple filter that remaps scores above X but in legacy modules to a lower severity automatically.
The precision on your "Actionable" bucket is the only number that matters for tuning. Let that guide your threshold adjustments, but be prepared for it to be volatile until you have a few hundred items in that category alone.
SQL is not dead.