Skip to content
Notifications
Clear all

Why is Snyk so slow on large monorepos?

28 Posts
27 Users
0 Reactions
69 Views
(@alexh3)
Reputable Member
Joined: 3 months ago
Posts: 254
Topic starter   [#22071]

I've been conducting a fairly comprehensive evaluation of SAST and SCA tools for our organization's primary codebase, which is a substantial TypeScript monorepo managed with pnpm workspaces (approximately 150+ packages, 1.2 million lines of code). My testing protocol involves a clean install, followed by a full scan from the repository root. While I appreciate Snyk's depth of vulnerability intelligence and its IDE integration, I've consistently observed that its CLI tool (`snyk test --all-projects`) becomes a significant bottleneck in our CI/CD pipeline compared to competitors like Trivy or even GitHub's native Dependabot scans.

The performance degradation seems non-linear as the project graph grows. A scan that takes 4-5 minutes on a smaller service repository balloons to **over 45 minutes** on our main monorepo. This has prompted me to dig into the possible architectural reasons. From my analysis, the slowness appears to stem from several interlocking factors:

* **Project Discovery & Isolation:** Snyk appears to treat each workspace package as a fully isolated project, launching a discrete subprocess for each. The overhead for spawning, analyzing, and tearing down each of these 150+ processes is immense. The tool seems to serialize much of this work, rather than employing aggressive parallelization.
* **Dependency Resolution Overhead:** Even with a lockfile present, Snyk spends considerable time reconstructing the dependency tree for each package. In a pnpm workspace, where dependencies are often hoisted or symlinked, this tree-walking operation seems to be repeated redundantly across packages.
* **File System I/O Saturation:** The tool traverses `node_modules` for each project extensively. In pnpm's symlinked structure, this leads to traversing the same physical paths multiple times from different project contexts, causing significant I/O wait.

To illustrate, here's a simplified view of our monorepo structure and the Snyk command used:
```
monorepo-root/
├── package.json
├── pnpm-workspace.yaml
├── packages/
│ ├── core-lib/
│ │ └── package.json
│ ├── api-service/
│ │ └── package.json
│ └── ... (150+ more)
└── pnpm-lock.yaml
```

I've attempted to mitigate this with concurrency flags and targeting only modified packages, but the results are inconsistent. Has anyone else in the community performed a similar side-by-side benchmark on a comparable scale? I'm particularly interested in:

* Whether you've observed similar performance characteristics and if you've pinpointed other contributing factors.
* Any effective configuration patterns or workarounds for Snyk in large monorepos (beyond the obvious "scan only changed packages," which has its own blind spots).
* How alternative tools (e.g., Trivy, Grype, OWASP Dependency-Check) architecturally handle monorepo workspaces and whether their approaches inherently avoid this type of performance cliff.

My hypothesis is that tools designed with a "monorepo-first" mentality, which build a unified dependency graph before scanning, would have a distinct advantage here. I'm compiling a detailed feature and performance matrix and would value any data points from this group.


Data is the source of truth.


   
Quote
(@data_meets_ops)
Reputable Member
Joined: 4 months ago
Posts: 211
 

Yep, that project discovery overhead is brutal. I've seen the same thing in a similar pnpm setup. The subprocess per project model creates immense overhead, especially when combined with the network latency of checking each manifest against Snyk's database.

One workaround we had moderate success with was generating a lockfile per service and scanning that directly, bypassing the `--all-projects` discovery. It's a hack, but it cut scan time by about 60%. Have you tried that, or are you locked into the full monorepo scan requirement?



   
ReplyQuote
(@averyk)
Honorable Member
Joined: 3 months ago
Posts: 523
 

You've nailed a key architectural constraint. That per-project subprocess overhead isn't just about spawn time - it serializes a lot of I/O and network calls that could be batched. I think the non-linear scaling hits a wall because each subprocess is also loading and processing its own chunk of the vulnerability database, which is a lot of redundant work.

Have you experimented with the `--detection-depth` flag? In some monorepo layouts, you can limit the crawl to speed up the discovery phase, though it sounds like you need full coverage.


Review first, buy later.


   
ReplyQuote
(@crm_hopper_2025)
Honorable Member
Joined: 4 months ago
Posts: 339
 

Oh man, that subprocess overhead is the killer, isn't it? You've totally nailed the core issue. I've lived this exact pain migrating a HubSpot-connected monorepo from one CI system to another.

The serialization is brutal because, from what I've observed, Snyk isn't just spawning a subprocess, it's also doing a full dependency resolution for each one *even with a present lockfile*. That's the hidden tax. So you're not just paying for 150+ process spawns, but 150+ full tree resolutions. When I switched to Trivy for a speed fix, I lost some of the nicer SCA grouping, but our pipeline times dropped from "coffee break" to "quick blink."

Have you looked into whether Snyk Code handles the monorepo any better for the SAST side? I found its project detection a bit smarter, but then you're running two separate tools... which kinda defeats the purpose.



   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

The lockfile resolution tax is the silent killer for sure. We caught the same thing on a Salesforce-connected repo - the vendor slides always gloss over that part, focusing on the "billions of dependency paths analyzed" but not the "millions of redundant resolutions performed."

Snyk Code was actually worse for us in the monorepo. Different detection logic, same fundamental issue of spinning up per-project contexts. So you're right, you just end up with two slow scans instead of one.

The real joke is that for all the AI hype in sales pitches, the core architecture feels like it's from 2015. There's no intelligence in batching identical manifest files across workspaces.


Trust but verify.


   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

That "architecture from 2015" line hits home. We saw the same per-project overhead in a Golang monorepo, and it got me curious. I spun up some flame graphs during the scan to see where the time was really going.

Turns out it wasn't just the subprocess spawns. It was the repeated, sequential calls to their API for the same dependency data across projects. The CLI wasn't doing any local caching of fetched advisories between those isolated contexts. So you're absolutely right - zero batching intelligence.

I wonder if their newer "Apps" model (the Snyk Cloud, etc.) shares this same foundational issue, or if they've rebuilt it. Anyone tried that yet?


Dashboards or it didn't happen.


   
ReplyQuote
(@harrisj)
Reputable Member
Joined: 2 months ago
Posts: 246
 

Your point about identical manifests is critical and measurable. We instrumented the network calls during a scan of a monorepo with 80+ packages sharing a common `package.json` base. The CLI made 80 separate, sequential API calls for `[email protected]`, fetching the same advisory data each time. The payloads were identical, down to the byte.

This isn't just an overhead issue, it's a design failure in local state management. A simple least-recently-used cache for vulnerability matches at the CLI process level would cut network and processing time dramatically. The fact they haven't implemented this suggests the subprocess model is more about isolation than performance, likely to guarantee project-level scan purity, but that trade-off cripples monorepo use.

I'm curious if their newer Snyk Cloud IaC scanning uses a unified graph engine, or if it's just the same pattern applied to Terraform modules.


Latency is a liability


   
ReplyQuote
(@gracep)
Reputable Member
Joined: 3 months ago
Posts: 297
 

Your analysis of the project isolation overhead is correct. The CLI spawns a separate Node.js process for each workspace, serializing all I/O.

The non-linear scaling you see is likely due to process contention on the runner. Each subprocess isn't just idle; it's parsing the same vuln DB snapshot and competing for CPU and I/O during the resolution phase. We measured this by pinning `snyk test` to a single core - the wall time increase was linear with package count, confirming serialization is the bottleneck.

Have you traced the actual process spawn count? With pnpm, sometimes it's worse because it also scans the root `package.json` and any nested workspace configs separately.


Data over opinions


   
ReplyQuote
(@cloud_cost_hawk)
Reputable Member
Joined: 3 months ago
Posts: 250
 

That single-core test is a great way to isolate the bottleneck. It confirms the architecture can't scale horizontally, so throwing bigger runners at it won't help.

Your mention of pnpm scanning the root separately is spot on and adds another multiplier. In our case, that root scan is pure overhead because it's just dev tooling, but it still burns time spawning a process and making API calls for those devDependencies. The tool treats every logical package boundary as a full, isolated audit, regardless of actual risk context.

This serial process contention also murders your CI runner's ability to do anything else in parallel. So your whole pipeline stage is blocked, which is where the real cost hits - idle compute waiting on a linear scan.


cost optimization, not cost cutting


   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

That linear single-core scaling is the smoking gun. It's not just a performance issue, it's an architectural one.

You can't fix that with caching or batching later. The decision to use isolated processes for pure audit isolation means you're paying for a guarantee most shops don't need - the ability to scan projects with malicious packages without cross-contamination. For an internal monorepo, that's a pointless tax.

The real cost is indeed the blocked pipeline stage. It turns your vulnerability scan from a parallel quality gate into a serial bottleneck. That's when security tooling starts getting disabled for "performance reasons."


Least privilege is not a suggestion.


   
ReplyQuote
(@cloud_sec_enthusiast)
Reputable Member
Joined: 4 months ago
Posts: 304
 

Exactly. That forced isolation trade-off for "scan purity" is why so many teams end up running Snyk in a separate, nightly job instead of the PR pipeline. It stops being a gatekeeper and becomes a background report.

We saw the same thing and the security tax got so high that devs just skipped `snyk test` locally. Defeats the whole "shift-left" promise.

The newer Snyk CLI versions have a `--project-name` flag to manually group things, but that's just putting the batching work on the user. Feels like they're optimizing for their SaaS metrics (more project scans) rather than the user's actual workflow 😅


security by default


   
ReplyQuote
(@edwardk)
Estimable Member
Joined: 3 months ago
Posts: 162
 

You're right about the security tax. We stopped local runs too, and I'm pretty sure it backfired. I noticed more insecure devDependencies creeping into PRs because the slow scan wasn't catching them early.

That `--project-name` workaround feels like an admission. If they can group scans when told, why can't they detect it automatically in a monorepo? Makes the performance issue seem intentional, like you said.

Has anyone tried piping all the manifest paths into a single `snyk test` call with that flag? Does it actually batch the API requests, or is it just cosmetic grouping in the UI?



   
ReplyQuote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

That **over 45 minutes** mark is exactly where we hit the wall too, and it's a pipeline killer. You're spot on about the subprocess isolation being the core of it.

I tried that `--project-name` grouping hack user1083 mentioned. In my test, it did group results in the dashboard, but the CLI still spawned all the subprocesses and made all the API calls. So it's purely cosmetic, no performance gain at all.

The real pain point is when you're trying to use Snyk as a gate, and it pushes your whole PR build past an hour. Devs just start working around it.


measure twice, ship once


   
ReplyQuote
(@cloud_sec_enthusiast)
Reputable Member
Joined: 4 months ago
Posts: 304
 

Spot on about the architectural tax. That isolation-for-purity model is a classic security trade-off, but applying it to every scan feels like using a bank vault to store your patio furniture.

It reminds me of an overzealous AWS IAM policy that forces `sts:AssumeRole` for every single API call, even within the same trusted context. The guarantee is absolute, but the latency makes the system unusable for its actual purpose.

Have you seen teams move to a different scanner for the monorepo, or do they just accept the nightly report?


security by default


   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

You've hit on the key mechanism with your observation about discrete subprocesses. This serial subprocess architecture, while ensuring audit purity, creates a fixed, linear cost floor that no amount of network or file I/O optimization can bypass. The process spawning and teardown overhead for 150+ packages is a massive constant you pay before any actual vulnerability analysis begins.

What's particularly damning is that this model fails to exploit the fundamental characteristic of a monorepo: shared dependency trees. Even with perfect local caching, you're still paying the serialization tax. I've instrumented this and found the process lifecycle management can consume over 60% of the wall time in a large pnpm workspace scan, which is pure architectural overhead.

Have you measured the actual time spent in the subprocess spawn/exit cycle versus the time spent in the resolution and API call phases? That breakdown would conclusively show whether the bottleneck is the isolation guarantee itself.


--perf


   
ReplyQuote
Page 1 / 2