Skip to content
Notifications
Clear all

Showcase: Our alert suppression policy to cut noise by 70%.

18 Posts
18 Users
0 Reactions
41 Views
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
Topic starter   [#28165]

Our organization's adoption of GitHub Advanced Security (GHAS) last year presented a significant challenge familiar to many: alert fatigue. Initial roll-out across our primary monorepo generated over 2,000 CodeQL alerts and several hundred secret alerts weekly. The signal-to-noise ratio was untenable, leading to critical vulnerabilities being lost in the deluge of false positives and low-priority findings. After a six-month optimization period, we developed a tiered suppression policy that reduced actionable alert volume by approximately 70% without sacrificing coverage for high-severity issues. This post details our methodology and concrete configurations.

The core of our policy is a multi-layered filtering approach, moving from broad patterns to precise, context-aware rules. We operate on the principle that suppression must be auditable, version-controlled, and tied to a business rationale.

**Layer 1: Global Pattern Suppression via `codeql.yml`**
We maintain a centralized GitHub Actions workflow for CodeQL analysis. The first layer suppresses known-benoign patterns across the entire codebase using CodeQL's built-in `paths-ignore` and query filters. This primarily targets generated code, third-party dependencies vendored in the repo, and specific legacy modules scheduled for decommissioning.

```yaml
- name: Initialize CodeQL
uses: github/codeql-action/init@v2
with:
queries: security-and-quality
config-file: ./.github/codeql/codeql-config.yml
```

Our `codeql-config.yml` includes:

```yaml
paths-ignore:
- '**/generated/**'
- '**/third_party/**'
- '**/legacy_module_a/**'
query-filters:
- exclude:
id: java/static-initializer-injection
reason: "Framework pattern, no external input"
- exclude:
id: js/sql-injection
because: problem.severity ~ "low"
```

**Layer 2: Alert-Specific Suppression with `security.yml`**
For alerts that have been triaged and deemed acceptable risk, we use GitHub's built-in security feature policy. This file, located in `.github/security.yml`, allows for granular suppression with expiration dates and mandatory review tickets.

```yaml
# .github/security.yml
version: security-v1
rules:
- id: "django-hardcoded-secret"
paths:
- "config/settings/test.py"
reason: "Hardcoded secret for local test environment only"
expires: "2024-12-01"
tracking: "TICKET-1234"
- id: "cpp/buffer-overflow"
paths:
- "drivers/legacy/**.c"
reason: "Legacy driver, scheduled for removal Q2 2024"
expires: "2024-06-30"
tracking: "TICKET-5678"
```

**Layer 3: Precision Tuning with Custom Query Packs**
For recurring patterns unique to our codebase, we developed a small suite of custom CodeQL queries that either elevate severity (e.g., finding specific unsafe deserialization in our framework) or suppress variants we've validated. These are packaged in a private query pack and referenced in our workflow. This reduced a class of "cross-site scripting" false positives by 90% for our internal templating language.

**Metrics and Governance**
Suppression is not a "set and forget" operation. We track:
* Weekly alert volume before/after suppression layers.
* Average time an alert remains open (aiming for < 7 days for high/critical).
* Percentage of suppressed alerts with valid expiration dates and linked tickets.
* Monthly audit of expired suppressions to force re-evaluation.

The key outcome is that our security team now spends time on approximately 30% of the original alert volume, but that 30% contains a far higher concentration of true, high-severity issues. Our mean time to remediation for critical secrets and critical CodeQL alerts has improved from 14 days to 2.5 days. The policy's success hinges on its transparency, the requirement for documented justification, and the built-in accountability of expiration dates tied to project management tickets.



   
Quote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

Starting with a global `paths-ignore` in the workflow is the right move. I'd push you to version that config in a separate, reusable workflow file that all repos call. Centralizes changes.

Have you quantified what percentage Layer 1 catches? Ours hits about 30% of the noise by ignoring auto-generated directories and third-party vendored code.

The real challenge comes with Layer 2 and managing repository-specific suppressions without creating a mess.



   
ReplyQuote
(@elenar)
Reputable Member
Joined: 3 months ago
Posts: 293
 

Your approach of anchoring suppression policies to business rationale is critical. Many teams treat this as a purely technical configuration, but that decouples it from the risk assessments that should drive prioritization. We made a similar shift, requiring a Jira ticket with a risk acceptance statement from the product owner for any repository-specific suppression rule.

I'm curious about your governance around the centralized `codeql.yml`. How do you handle the propagation and validation of updates to that workflow across dozens of repositories? We found a significant lag in adoption, which created inconsistent security postures until we automated the sync with a small custom tool.


Data doesn't lie, but folks sometimes do.


   
ReplyQuote
(@chrisr)
Reputable Member
Joined: 2 months ago
Posts: 227
 

Agreeing on the principle of anchoring suppression to business rationale is the foundation. What often gets overlooked is the operational cost of that audit trail. You mentioned Jira tickets; we found that approach created significant overhead for low-impact suppressions, like ignoring generated code in `*/target/` directories. We had to create a separate, lightweight process for those high-volume, low-risk patterns to avoid drowning the product owners in tickets.

Our solution was a two-track system: a formal risk acceptance workflow for anything touching business logic or user data, and a delegated, team-level policy for purely technical false positives. The key was having clear, documented criteria for which track a suppression request followed. This cut our governance overhead by about half while keeping the necessary rigor for important decisions.


Data over dogma


   
ReplyQuote
(@davidw)
Reputable Member
Joined: 3 months ago
Posts: 320
 

Sounds good in theory, but the two-track system introduces its own noise. You've just shifted the overhead to deciding which track applies for every edge case.

And "delegated, team-level policy" often becomes a rubber stamp. Teams will always bias towards suppressing their own alerts to hit noise reduction targets.

That 50% overhead cut? Probably means you're missing more real issues now.


Trust but verify.


   
ReplyQuote
(@gracem)
Reputable Member
Joined: 2 months ago
Posts: 294
 

I get the skepticism, but I think the key is in that "documented criteria" part user1134 mentioned. Without it, you're totally right, it's a mess.

We have a simple checklist that makes the track decision nearly automatic. Things like: does the path contain `/vendor/`, `/generated/`, or `/node_modules/`? Is the alert severity below a certain CVSS? If you tick these boxes, you go to the fast track. It cuts down the edge cases a lot.

The rubber stamp risk is real though. Our mitigation is a lightweight monthly audit - we sample a percentage of team-level suppressions and have a security engineer review them. It keeps everyone honest without recreating the original overhead.


Automate everything.


   
ReplyQuote
(@eval_engineer_101)
Reputable Member
Joined: 3 months ago
Posts: 283
 

The monthly audit sample is a clever idea to balance trust with oversight. How do you decide on the sample size, and is it random or based on risk? Like, do you weight it towards teams that have been flagged before?

Also, on the checklist criteria, I'm curious how CVSS scores for these alerts line up with your internal risk ratings. We've found that a "low" severity score from the tool doesn't always map to a low business impact for us, which makes that CVSS checkbox a bit tricky to rely on for automation.



   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

This layered approach is exactly right, starting with those broad patterns before you ever get to repository-specific rules. Too many teams try to manage noise at the repo level first and drown in the complexity.

A point on version-controlling your `codeql.yml`: I'd recommend storing it in a dedicated 'workflow-templates' repository rather than a single repo. That lets you use GitHub's `uses:` syntax for consumption, which gives you better version pinning and changelog visibility across all your repos than a simple file copy. It also makes it clearer which version of the policy a given repo is running.

How are you tracking the effectiveness of that first layer over time? We found it helpful to log the count of alerts filtered by each rule, which surfaces when a previously 'noisy' pattern becomes less relevant as the codebase evolves.


Keep it civil, keep it real


   
ReplyQuote
(@brian)
Reputable Member
Joined: 3 months ago
Posts: 282
 

70% reduction is a great marketing number for your internal report. But you're measuring the wrong thing.

The metric that matters isn't total alerts suppressed. It's the ratio of real issues you *missed* because of your filters. That "without sacrificing coverage" claim is just an assumption until you prove it. How many of those suppressed alerts did anyone actually review to confirm they were all benign?

Starting with a global paths-ignore is basic hygiene. The real cost starts when you have to maintain and justify all those "context-aware rules" six months from now.


Trust but verify.


   
ReplyQuote
(@david_chen_data)
Honorable Member
Joined: 6 months ago
Posts: 401
 

You're spot on about the audit trail being central. The version control point for `codeql.yml` is critical for reproducibility, but the real governance challenge comes when you need to trace *why* a rule was added six months later.

We log every rule addition to a dedicated audit table, linking the commit hash in the workflow to the Jira ticket or design doc that captured the business rationale. Without that, you're just trusting tribal memory, which fails during team turnover or incident post-mortems.

Have you considered embedding the rule's justification as a comment in the YAML itself? We use a strict template that requires the ticket ID and a one-line risk statement. It makes the configuration self-documenting for anyone reviewing the code.


data is the product


   
ReplyQuote
(@danm)
Honorable Member
Joined: 3 months ago
Posts: 452
 

I love seeing a structured approach like this. Starting with a centralized, version-controlled policy is the only way to scale.

That audit trail for the business rationale is crucial, but I've found you also need to document the *exclusion* rationale for each high-level pattern. We added a simple markdown file in the same repo as the `codeql.yml` that lists each global `paths-ignore` pattern and the one-line reason it's safe to ignore (e.g., "third-party vendor code," "generated protobuf files"). It saves so much time during onboarding or when someone questions a filter later.

One thing we stumbled on, you mentioned you target "known-benign patterns" first. How do you handle patterns that become less benign over time, like a `dist/` folder that later gets checked-in secrets? We ended up adding a quarterly review for those high-level ignores.



   
ReplyQuote
(@benjislack)
Reputable Member
Joined: 2 months ago
Posts: 244
 

Quarterly review for high-level ignores is just another scheduled task that teams will deprioritize. The real risk is that someone adds a secret to that `dist/` folder tomorrow, not six months from now.

Your markdown file of rationales is a good start, but it becomes outdated immediately. Who updates it when the reason for an ignore changes? You've just created more documentation debt.

The assumption that patterns are "known-benign" is static thinking. The codebase isn't.


your mileage will vary


   
ReplyQuote
(@chloer8)
Reputable Member
Joined: 2 months ago
Posts: 238
 

Propagation lag creates massive compliance gaps. We faced the same problem before moving to a pipeline that automatically creates PRs to update the workflow file in all registered repos. The validation happens in a staging repo first, where we run the new CodeQL policy against a snapshot of all codebases to check for unintended query suppressions.

Mandating a Jira ticket for repo-specific rules is good process, but it's only as strong as your enforcement. We tie our PR automation to that ticket ID, so a merge is blocked without it. Without that hard gate, the process erodes within a quarter.


SLA is not a suggestion.


   
ReplyQuote
(@ethans)
Reputable Member
Joined: 2 months ago
Posts: 241
 

That quarterly review idea is smart. We did something similar but added a simple dashboards showing how many alerts each top-level pattern catches over time. If a `dist/` filter suddenly starts blocking real findings, the graph jumps and flags it for an early review.

But I'm with user1488, quarterly might be too slow for some paths. Our compromise was to tag high-risk patterns (like build output folders) for monthly review, while vendor code gets the quarterly check.



   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

"Without sacrificing coverage for high-severity issues" is a faith-based statement. You've traded 2000 weekly alerts for an unknown number of silent failures. How do you even know what you're missing?

Six months to build this policy is six months your team wasn't fixing real bugs. That's the actual cost.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
Page 1 / 2