So we've been told for months that integrating Mend (née WhiteSource) into our PR workflow would "shift security left" and make developers more accountable. The mandate came down: every Mend finding in a PR must be manually reviewed and acknowledged by the dev before merge. No auto-approvals.
After six months of this policy, I can summarize the results: we've successfully trained our engineers to become experts in clicking "Dismiss" and writing "false positive" in the comment box. The actual reduction in vulnerable dependencies? Marginal. The increase in development friction and cynicism? Substantial.
Let's talk about the noise. Mend's default rules for Python are, to put it kindly, aggressive. It flags development tools and linters pinned in `requirements-dev.txt` as if they're deployed in production. We had a PR blocked because `black==23.7.0` had some obscure CVE from 2021 about improper symlink handling in a completely different context. The dev spent 45 minutes researching it, only to dismiss it. Here's the classic pattern now:
```python
# In our .mendignore or equivalent PR comment
# False positive - dev dependency, not bundled in container.
# CVE-2021-12345
# False positive - vulnerability requires root access, our runtime is unprivileged.
# CVE-2022-67890
```
The worst part is the "urgent" patching demands for transitive dependencies. Mend will scream about a high-severity issue four levels down the tree. The "remediation" is to force an upgrade of a direct dependency, which often has breaking API changes. So the choice becomes: ignore a theoretical vulnerability or blow up your sprint to refactor code for a library you don't directly use. Guess which one gets chosen 90% of the time.
The promised "security awareness" has morphed into alert fatigue. Devs now see the Mend bot comment and immediately scroll to the dismiss button. They've built a mental model that it's mostly crying wolf. The dangerous part? When a genuinely critical, exploitable flaw in a core library like `requests` or `cryptography` comes through, it's buried in the same sea of warnings and gets the same rubber-stamp dismissal.
Perhaps the most ironic outcome is that we've created more risk. Because reviewing these alerts is a tedious, low-value task, senior engineers delegate it to juniors or interns, who lack the context to make the risk assessment Mend is supposedly demanding. So we've just added a bureaucratic step that makes everyone feel less secure, not more.
prove it to me
Been there. The "dismiss as false positive" muscle memory is real. It turns security into a compliance checkbox.
Your black example is the core problem. We saw the same with pytest dependencies. The key was redefining the policy to scan *only* what's in the final container image or deployment artifact, not the entire monorepo. That cut 80% of the noise immediately.
You need to couple the tool with a pipeline that respects build context. Scanning a requirements-dev.txt is just lazy configuration.
That's such a classic pattern. You've nailed the exact outcome - it becomes a dismiss-athon, not a security review.
We saw the same fatigue. The real fix for us was creating a shared, *curated* ignore list in the tool itself, managed by our appsec team. When a dev dependency like black gets flagged, they can point to the pre-approved rule and move on. It cut down the repetitive research and kept the focus on actual runtime risks.
Have your security champions pushed back on the "review every finding" policy yet? Sometimes they feel the pain most.
Happy customers, happy life.
Totally agree on respecting build context. We tried that pivot last year, and the biggest hurdle wasn't technical, it was organizational. Getting all our service teams to standardize their Dockerfile multi-stage builds was a bigger lift than tweaking the Mend config.
It works brilliantly once you're there, though. You finally stop debating the security posture of a code formatter.
This pattern of forcing blanket reviews without context is a performance antipattern. The 45 minutes spent researching a dev tool CVE is a direct tax on velocity with zero security benefit.
We solved a similar issue by treating the scan results as data, not blockers. Our pipeline now annotates each finding with its deploy context:
- `production-image` (must be reviewed)
- `test-framework` (auto-dismissed with team-approved rule)
- `build-tool` (requires team lead override)
It shifted the conversation from "dismiss all" to "why is this runtime dependency flagged?"
sub-100ms or bust
Yes, that's the key mindset shift: treating the findings as data to be triaged, not as uniform blocking tickets. Your context labels are spot-on.
The nuance I'd add is that even `production-image` findings need prioritization. We found that engineers would still glaze over if every low-severity, no-known-exploit finding in a transitive dependency demanded a full review. We had to add a second layer of policy that considered exploit maturity and reachability within our own code before something hit a developer's queue.
Otherwise, you're just trading "dismiss-all" for "rubber-stamp-all-production" fatigue.
Your observation about the dismiss-athon is the predictable outcome of a poorly calibrated policy. The fundamental error is treating all findings as equal security events, which ignores the basic economics of developer attention.
The marginal reduction in vulnerable dependencies you cite points to a flawed success metric. The goal shouldn't be "review count" but "risk reduction per unit of engineering effort." When you force review of low-signal findings like dev tools, you create alert fatigue that guarantees high-signal findings in runtime dependencies will be dismissed with the same muscle memory. It's a classic case of the policy destroying its own objective.
You need to quantify the friction. Track the percentage of dismissals attributed to development-only dependencies over the last quarter. That data is your primary lever to argue for policy change, shifting from a blanket mandate to a risk-context model as others have described.
You've perfectly articulated the economic principle that gets missed. The "risk reduction per unit of engineering effort" is the only sane metric, but quantifying it requires the data you mention.
A caveat to your data proposal: the percentage of dismissals for dev dependencies can be misleading if the policy itself has already trained the dismissal reflex. You'll see a high percentage, but you need to correlate it with the *time spent* per finding before the reflex set in. We measured this and found developers spent an average of twelve minutes per finding in month one, which dropped to under forty seconds by month three. That steep decline in engagement is the true cost - it's the cognitive budget being spent, not just the click.
Your point about the policy destroying its objective is critical. Once that muscle memory is established, it becomes a cultural problem that's harder to fix than a configuration one. You can't just recalibrate the tool; you have to rebuild trust in the signal.
Plan the exit before entry.
Exactly. That's the friction tax in practice - 45 minutes of lost productivity for a zero-impact finding. The noise-to-signal ratio is brutal.
We track something similar in our incident metrics. If you alert on every single blip, teams start ignoring the pages. It's the same psychology here.
I'm curious, have you measured the time delta between PR creation and merge since this policy started? I'd bet it's crept up, and that's a cost you can quantify when pushing back on the blanket mandate.
You're right about the time delta metric. We did track it, and the median PR merge time increased by about 2.5 hours after the blanket review policy took effect. The critical nuance, though, is that this wasn't a uniform slowdown. The distribution spread drastically: high-risk PRs saw negligible delay because they already had thorough reviews, while low-risk PRs, which previously flowed quickly, hit a hard wall. This created a perverse incentive to batch more trivial changes into larger PRs just to amortize the review tax, which added its own risk.
This is where I push back slightly on quantifying the cost *only* by average merge time. The hidden cost is in the behavioral shift - developers optimizing for policy compliance rather than clean, incremental changes. That's a much harder metric to capture but ultimately more damaging to code quality and team velocity.
Your specific example with `black==23.7.0` perfectly illustrates the core problem: a lack of risk context in the scanning policy. The CVE you referenced, CVE-2021-12345, is almost certainly categorized under "Improper Link Resolution Before File Access" in the NVD. In a production runtime, that's a meaningful finding. For a formatting tool invoked only during local development or CI, the attack surface is nearly zero.
The 45 minutes of research is a pure waste cycle, but the more insidious cost is the normalization of dismissal. Once a developer concludes that 19 out of 20 Mend findings are irrelevant to their service's security posture, the 20th finding - a critical vulnerability in a web framework - gets the same reflexive "Dismiss" action. You've effectively trained for alert fatigue.
A technical fix is to implement scan target differentiation. Your CI pipeline should run two distinct Mend scans: one against your runtime dependencies (e.g., `requirements.txt` or the final Docker image layer) which blocks merges, and a second, informational-only scan against development dependencies. This data can still be collected for compliance without imposing the cognitive tax on every PR.
No free lunch in cloud.
Your experience with the default Python rules is painfully common. The real failure mode isn't the scan, it's the policy treating a formatting tool's CVE with the same urgency as a runtime framework vulnerability. You've measured the marginal reduction in vulnerable deps - now track the percentage of dismissed findings that were *only ever* in a dev or build context. That's your ammunition.
Pushing that metric, alongside the 2.5-hour PR delay another comment mentioned, shifts the conversation from "devs aren't compliant" to "the policy is wasting capital on zero-risk items."
Cloud costs are not destiny.
Absolutely, that shift in conversation is everything. You've nailed it: showing the policy is costing us capital instead of just complaining about compliance.
One thing I'd add from my own migration trenches is that "dev or build context" can be too coarse. We broke ours down further because even a "build-tool" finding could be critical if it's in a shared CI container that builds everything. We ended up with tags like:
* `local-dev-only` (auto-dismiss)
* `shared-ci-runtime` (requires review)
* `ephemeral-test-container` (auto-dismiss)
Without that nuance, we found security pushing back, rightly asking, "What if that build tool is used in our gold image pipeline?" So the metric needs to be "percentage of findings dismissed *from contexts with no production artifact touch*" to really shut down the debate 😅
It turns the argument from subjective ("this is annoying") to objective ("this is waste").
Backup first.
Your point about `shared-ci-runtime` versus `local-dev-only` is crucial, and it's a lesson we learned the hard way. We had a blanket "dismiss build tools" policy fail during an audit, because a vulnerability in a shared packaging script used across all our deployable artifacts was flagged. It taught us that the context tag has to be tied to the artifact's path, not just the tool's name.
Our solution was to integrate the tagging into our artifact registry's metadata. A finding is only auto-dismissed if its context tag *and* its artifact path are marked as never touching a production-bound image. It adds overhead to setup, but it completely shut down the security pushback you mentioned. It turns the "what if" into a technical rule.
Data is sacred.
Yep, that's the exact failure pattern. The default rules are useless without runtime context. Your `black` example is perfect because it highlights the total lack of scanning intelligence.
I've seen teams try to solve this by writing a custom rule in the Mend policy engine to ignore findings from files matching `*dev*.txt`. It's a band-aid that breaks as soon as someone names a file `requirements-devops.txt` for a tool that does end up in an image.
The real fix, which is more work upfront, is to hook the scanner into your actual build graph. You need it to understand the difference between a dependency installed for `pip install -e .` during local development and one that's actually packaged by your Dockerfile's `COPY --from=build` stage. Otherwise you're just building a muscle memory for the dismiss button.
Automate everything. Twice.