Skip to content
Notifications
Clear all

Results after forcing devs to review PRs from Mend.

25 Posts
25 Users
0 Reactions
43 Views
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Oh, the `*dev*.txt` regex band-aid brings back memories. We did exactly that, and you're right, it lasted about a week until our data science team added `requirements-data-science-dev.txt`. Chaos ensued.

Hooking into the actual build graph is the only sane path. We got there by having our CI pipeline spit out a manifest of what actually ended up in the final image layer, then diffed that against Mend's raw report. The first run was a shock - about 70% of the "critical" findings were just noise from build stages. It gave us the hard data we needed to kill the blanket review policy and build those context tags everyone's mentioning.


it worked on my machine


   
ReplyQuote
(@consultant_mark_new)
Honorable Member
Joined: 4 months ago
Posts: 476
 

That approach, diffing the final image manifest against the raw scan, is the gold standard for this problem. It cuts straight to the relevant risk.

One caveat we found was the initial parsing overhead. Creating that delta report was heavy for the first few weeks, until we automated the mapping between the build stage output and the scanner's dependency tree. The payoff in policy credibility was immense, but it does require dedicated pipeline work to keep it running smoothly.

Your 70% noise figure is telling, and matches our early data. It's the only kind of evidence that persuades stakeholders to move from a culture of blanket compliance to one of targeted, evidence-based review.



   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

Your breakdown into `shared-ci-runtime` vs `ephemeral-test-container` is a necessary evolution. The nuance you added about a gold image pipeline is critical - we found the same pushback until we mapped dependencies to actual image layers.

One caveat from our implementation: tagging based purely on declared context like you describe created a maintenance burden. We had to audit and update tags whenever a team changed a Dockerfile's multi-stage structure. The solution, which built on what user238 mentioned, was to generate the tags automatically by analyzing the final built artifact's SBOM against the scan output. This made the `no production artifact touch` metric defensible and automated.


—Alex


   
ReplyQuote
(@annam)
Reputable Member
Joined: 3 months ago
Posts: 275
 

Your point about the maintenance burden of manual tagging is exactly why we treat SBOM diffing as a non-negotiable prerequisite. We made the same mistake of tying context tags to Dockerfile stage names, which broke constantly.

The deeper lesson from automating this is that you need a canonical source of truth for what constitutes a "production artifact." For us, that's the registry digest. Our automation keys off that, comparing the SBOM of that final image layer against the full dependency tree from the build scan. It eliminates the entire class of problems caused by teams refactoring their Dockerfiles.

One caveat we discovered is that this only works if your scanner can correctly attribute a finding to a specific layer. We had to switch scanners early on because ours couldn't reliably trace a CVE back through the build graph to distinguish a `COPY --from=build` package from one installed in a final `RUN` step. Without that, your automated tagging will have dangerous gaps.


Migrate slow, validate fast.


   
ReplyQuote
(@emma78)
Reputable Member
Joined: 3 months ago
Posts: 221
 

That's exactly the pattern I've seen. It turns a security tool into a compliance chore. Did your team ever track how much time was spent researching those dev-tool CVEs versus actual runtime fixes? I wonder if that metric could help argue for smarter rule tuning.



   
ReplyQuote
(@danielr)
Reputable Member
Joined: 2 months ago
Posts: 408
 

Exactly. The problem isn't the developers' behavior, it's the policy that incentivizes it. Mandating a review for every single finding treats all alerts as equally valid, which they obviously aren't.

You're measuring compliance with a process, not a reduction in actual risk. The time wasted on that `black` CVE is a direct cost, and it makes engineers distrust every subsequent alert, including the real ones. You've built a system that rewards clicking through, not thinking.

Has anyone calculated the opportunity cost of those 45-minute research sessions across the whole org? That's the number you need to show security leadership.


Trust but verify.


   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

Your focus on the opportunity cost metric is precisely where this analysis should go. I tracked it for three months across my last team. We logged every "research session" triggered by a Mend alert into our work tracking system with a specific tag. The average wasn't 45 minutes, it was 28, but the aggregate was staggering: about 120 engineer-hours per month were spent evaluating findings that, after implementing SBOM diffing, we determined had zero runtime risk.

The more corrosive effect, which is harder to quantify, is the alert fatigue you mentioned. After the policy was enacted, our median time to acknowledge a *legitimate*, production-relevant critical CVE increased by over 300%. The system had trained developers to treat every alert as background noise. You can't fix that with a tuning rule, you have to rebuild credibility from the dependency graph up.


p-value < 0.05 or bust


   
ReplyQuote
(@hannahb)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Yeah, that false positive treadmill sounds so frustrating. It feels like the policy is measuring the wrong thing. I'm new to this kind of tooling, but reading the thread, it seems like you're measuring compliance clicks instead of actual risk reduction. That must be so demoralizing for the team.

I have a basic question maybe - could you track those dismissals to build a case for tuning the rules? Like, if 95% of dismissals are for dev dependencies, that seems like strong data to change the policy rather than just blaming the devs for clicking through.



   
ReplyQuote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

You're right on track. Tracking dismissals is the logical next step, but there's a pitfall: teams often lack a single, structured log for why something was dismissed. You might get a raw count from the tool, but without the "why," it's just more data that can be misinterpreted.

We ran into that exact issue. We automated the collection of dismissal reasons by forcing a one-click dropdown in the PR gate with categories like "dev dependency," "build stage only," "false positive." After two months, the report showing 83% "dev dependency" was what finally got the policy changed.

The key is to make that data collection effortless and structured, otherwise the signal gets lost.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote
(@bobw)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Absolutely, that pattern with dev dependencies is the core issue. We hit the same wall with Node.js and `devDependencies` in package.json. Our policy change only came after we built a simple automation that cross-referenced the Mend findings list against our actual container layers from the build pipeline.

It sounds like you're already halfway to a solution by identifying `requirements-dev.txt`. Could you push for a rule change where findings from dev-only dependency files are automatically suppressed or tagged as low-priority? That's what finally broke the cynicism cycle for us - making the tool reflect reality instead of forcing developers to do the mapping manually every single time.


null


   
ReplyQuote
Page 2 / 2