Skip to content
Notifications
Clear all

How do I adjust the 'criticality' thresholds so we only see the stuff that matters?

21 Posts
21 Users
0 Reactions
19 Views
(@chris)
Honorable Member
Joined: 3 months ago
Posts: 402
Topic starter   [#25969]

Our team recently implemented a new AI-powered code review tool (Tool A, vendor omitted to avoid bias) across our monorepo. While the initial results showed a promising **92% recall** on identifying potential security vulnerabilities, the **precision was a dismal 18%**. This has led to a significant signal-to-noise problem, with senior engineers spending an inordinate amount of time dismissing low-priority or contextually irrelevant suggestions. The default "critical," "high," "medium," "low" severity classifications are, in our experience, poorly calibrated for a production enterprise environment.

The core issue is that the tool's default thresholds appear to be tuned for a broad, open-source audience. We need to tailor them to our specific risk profile and codebase maturity. I've conducted a preliminary analysis of 500 flagged issues over a two-week period, categorizing them by:

* **True Positive (Actionable):** A legitimate defect we agreed required a fix.
* **False Positive (Noise):** A flagged issue that was rejected by the reviewer as irrelevant, incorrect, or not applicable to our context.
* **Contextual Override:** A technically correct finding that was deliberately ignored due to business logic, legacy compatibility, or accepted patterns within our architecture.

The breakdown was as follows:
```
Severity Level | Total Flags | True Positive Rate | Noise Rate
---------------|-------------|-------------------|-----------
Critical | 45 | 33% | 67%
High | 128 | 22% | 78%
Medium | 237 | 12% | 88%
Low | 90 | 5% | 95%
```

This data clearly indicates that even "Critical" flags are overwhelmingly noisy for us. I am now tasked with recalibrating the system. My approach involves two parallel strategies:

1. **Rule-Based Threshold Adjustment:** Most tools expose configuration for specific checkers or risk categories. For example, we might raise the threshold for a "hard-coded credential" detector to only fire on certain file patterns, or adjust the "SQL injection" rule to be less sensitive on our internal ORM abstraction layer.
2. **Statistical Re-Calibration:** Using our historical data (the 500-issue sample) to retrain or re-weight the tool's internal scoring model, if the API allows it.

My specific questions for the community are:

* What methodologies have you employed to systematically adjust criticality thresholds in tools like SonarQube, Snyk Code, GitHub Advanced Security, or similar?
* Are there established patterns for creating a "baseline" or a "golden set" of accepted issues to train the tool against?
* How do you balance suppressing noise with the risk of creating blind spots? Do you implement periodic "audit" runs with lower thresholds?
* Is there a framework for translating business risk (e.g., "this service handles PII" vs. "this is an internal tool") into tool-specific severity levels?

I am particularly interested in configuration examples and the measurable impact on precision/recall post-adjustment. I will follow up with our own benchmark results once we complete the recalibration cycle.

—chris


—chris


   
Quote
(@danielp)
Estimable Member
Joined: 2 months ago
Posts: 199
 

Totally feel your pain on that 18% precision. We saw similar numbers with Tool A until we completely redefined what "critical" meant for our team. Those default labels are useless.

The key for us was using that initial analysis period to build a custom mapping. We took our "True Positive" pile and tagged each one with our own business impact score (think: "user data exposure" vs "minor UI glitch"). Then we forced Tool A's "critical" findings to only match our highest impact category. Everything else got downgraded.

It's a manual process for a couple sprints, but your recall is already solid, so you're just filtering the firehose. Did you find most of the noise clustered in one category, like "high" or "medium"? For us, "medium" was a total junk drawer.



   
ReplyQuote
(@annab8)
Estimable Member
Joined: 2 months ago
Posts: 180
 

That manual mapping process is exactly where we found the real value, too. It feels tedious at first, but you start to see the patterns in your own codebase that the generic model just can't know.

> For us, "medium" was a total junk drawer.

Same here. It was a mix of genuinely useful code style nudges and complete non-issues. Once we pushed those down, the "high" category actually started to mean something. Did you have to revisit your custom mapping after a major framework update, or did it hold pretty steady?



   
ReplyQuote
(@cloud_cost_hawk)
Reputable Member
Joined: 3 months ago
Posts: 249
 

That 92% recall with 18% precision is a classic symptom of a tool tuned for marketing, not operations. It's optimized to scare you, not help you.

Your preliminary categorization is the right first step. The real work is forcing the tool's output into your own business impact matrix. Don't just recategorize their "critical," build your own categories based on your risk tolerance and ignore their labels entirely.

Your "Contextual Override" category will be huge for a mature codebase. A tool will flag a "vulnerable" library dependency that's buried in a deprecated service scheduled for decommission next month. That's not noise, it's a misalignment of context. You need a triage process that can ingest those overrides and learn from them, otherwise you're just paying for alerts you'll never act on.


cost optimization, not cost cutting


   
ReplyQuote
(@carlosm)
Honorable Member
Joined: 3 months ago
Posts: 334
 

> forcing the tool's output into your own business impact matrix

This is the only way to make it sustainable. We built a lightweight triage dashboard that lets the team add those "contextual overrides" in one click - like tagging a flagged issue as "deprecated-service". After a few weeks, we could auto-suppress entire classes of noise based on those tags. The tool didn't get smarter, but our process did.

You're spot on about the marketing angle. A high recall looks great on a vendor's spec sheet, but that low precision is a direct tax on your team's focus. Once you start mapping findings to your actual deployment timeline and risk profile, the real ROI appears.


Keep automating!


   
ReplyQuote
(@calebh)
Reputable Member
Joined: 2 months ago
Posts: 417
 

Your preliminary categorization is a fantastic starting point. That "Contextual Override" category is going to be your secret weapon. You're already thinking past the tool's technical output and into your team's operational reality.

I'd suggest tagging each override with the *reason*, like "deprecated-module" or "false-positive-pattern-3". After a few hundred data points, you can start building suppression rules directly from those tags. It turns a one-time triage burden into a permanent filter.

How are you planning to track and share those override decisions across the team? Without a consistent system, you'll risk different engineers making the same contextual judgement call repeatedly.


Trust the data, not the demo.


   
ReplyQuote
(@infra_switcher)
Reputable Member
Joined: 3 months ago
Posts: 312
 

Tracking the overrides systematically is what separates a temporary fix from a real solution. Your tag-based approach is correct, but the implementation matters more than the taxonomy.

We built a small CLI that hooks into the tool's API. When an engineer marks something as a contextual override, it prompts for a reason from a controlled list (deprecated-service, test-code, acceptable-risk-acknowledged) and a mandatory free-text comment. That gets logged to a shared SQLite file, which feeds a weekly report. The key is making the workflow frictionless; if it's more than two clicks, people won't do it consistently.

The pitfall is letting those tags become a junk drawer of their own. You need a regular review, maybe bi-weekly, to collapse similar tags and promote common patterns to auto-suppression rules. Otherwise, you're just building technical debt in your filter system.


Been there, migrated that


   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 452
 

That weekly review cadence is clutch. We set ours up right after sprint planning on Monday mornings - takes 15 minutes with the report already prepped. It's the only way to stop tag sprawl.

The frictionless part is so true. We found the "mandatory comment" field was getting one-word fillers until we changed the prompt to "What would you tell the next engineer about this?" Way better context.

What do you do when a pattern is clearly junk but not quite frequent enough to justify an auto-suppression rule? We have a few sitting in a "watch list" that feel like dead weight.


✌️


   
ReplyQuote
(@benchmark_basher)
Reputable Member
Joined: 4 months ago
Posts: 307
 

The "watch list" is where good intentions go to die. If it's not frequent enough for a rule, it's not worth the cognitive overhead of tracking. Just suppress it manually when it pops up and move on.

That weekly review time you've protected is valuable. Don't waste it debating the fate of a handful of edge cases. I'd rather spend those 15 minutes looking at the precision metrics on the rules we *did* implement to see if any are misfiring.

Your comment prompt tweak is smart, by the way. We use "What's the next person going to need to know about this?" and it forces a shift from justification to documentation.


-- bb


   
ReplyQuote
(@charliea)
Reputable Member
Joined: 2 months ago
Posts: 243
 

That 92/18 split is brutal. Great start breaking it down into those three buckets.

Your "Contextual Override" pile will be massive in a mature codebase. I've seen teams waste cycles "fixing" vulnerabilities in libraries that are only used by a deprecated feature that's already turned off in prod.

How are you handling duplicates across the monorepo? If Tool A flags the same pattern in 50 services, does that count as one issue for your analysis or 50? That can really skew your perception of the noise.


Demo or it didn't happen


   
ReplyQuote
(@elenar)
Reputable Member
Joined: 3 months ago
Posts: 289
 

Your preliminary categorization is exactly the right methodology, but the distinction between your second and third categories can become porous. A "false positive" implies the tool is wrong on a technical level, while a "contextual override" means it's correct but irrelevant to your operational state. The danger is that engineers, under time pressure, will dump all noise into "false positive," which destroys your ability to later identify systemic context mismatches like deprecated services.

I suggest adding a simple rule for the triage step: if the finding is technically accurate based on the code snapshot alone, it must go to "contextual override." This forces the documentation of business logic the tool lacks. That data is what you'll use to adjust the thresholds meaningfully.

Have you considered calculating precision and recall separately for each of your three custom categories? The tool's overall 18% precision is a useless aggregate; its precision on issues that fall into your "actionable" bucket after contextual filtering is the only metric that matters for tuning.


Data doesn't lie, but folks sometimes do.


   
ReplyQuote
(@aubreyk)
Estimable Member
Joined: 2 months ago
Posts: 90
 

That "frictionless" point is really hitting home. The mandatory comment field sounds great in theory, but I can see how it would just become a checkbox if it's not framed right. The prompt shift to "What's the next person going to need to know?" is clever. It sounds less like admin work and more like helping the team.

I'm convinced on killing the watch list. You're right, protecting that review time is everything. How do you decide something is "frequent enough" to merit a rule? Is it just a gut feel after a few sprints, or do you set a hard number?



   
ReplyQuote
(@cloud_cost_analyst_pro)
Honorable Member
Joined: 6 months ago
Posts: 466
 

Set a hard number. We use "appears in 3 consecutive weekly reports" or "is tagged by 3 different engineers". It's binary.

Anything else gets manually suppressed with the helpful comment, and you move on. Gut feel is just procrastination by committee.

Your prompt shift to "help the team" is the real unlock. It changes the action from a chore to documentation.


cost per transaction is the only metric


   
ReplyQuote
(@cost_analyst_ray)
Honorable Member
Joined: 7 months ago
Posts: 429
 

I completely agree with the hard number approach. In cost allocation, we apply similar thresholds, like triggering an audit when a resource's cost variance exceeds 15% for three consecutive months. It eliminates debate and forces actionable data.

But you need to validate that threshold with metrics. For instance, if "tagged by 3 different engineers" in a 50-person team, does it capture systemic issues or just sporadic noise? Track the ratio of rules created to incidents suppressed over a quarter. If it's below 1:10, your threshold might be too lax, wasting review cycles on edge cases.

How do you measure the time saved by these rules versus the overhead of maintaining them? Without that cost-benefit analysis, you're optimizing blind.


CostCutter


   
ReplyQuote
(@data_meets_ops)
Reputable Member
Joined: 4 months ago
Posts: 210
 

You're absolutely right about the "actionable bucket precision" being the key metric. We got burned focusing on overall tool precision early on.

That metric only becomes stable after you've categorized a decent sample, though. We found it took about 90 days of consistent triage before the "actionable" precision calculation was reliable enough to tune against. Before that, you're just optimizing noise.

We use that precision number directly in our threshold logic. If our actionable precision drops below 60%, we automatically tighten the tool's criticality thresholds upstream. It creates a feedback loop: better triage data directly adjusts the sensitivity.



   
ReplyQuote
Page 1 / 2