Trying to integrate Semgrep into our CI. The pitch was "fast, easy SAST." So far, it's just slow.
I want it to run *only* on the changed files in a pull request, like a linter. Everyone says to use `semgrep ci`, but that seems to require their proprietary cloud platform. We run everything on-prem.
I've tried passing a list of changed files to `semgrep --config auto .`. It still wants to scan the whole repo. The `--include` flag is a nightmare to get right.
What's the actual, usable pattern here? My current hack is generating a `.semgrepignore` with all unchanged files, which is ridiculous.
Example of what I'm wrestling with:
```bash
# Get changed .py files in PR
CHANGED_FILES=$(git diff --name-only origin/main... -- '*.py')
# How do I feed this to semgrep without scanning everything?
semgrep --config auto $CHANGED_FILES # Often fails on paths, misses dependencies
```
If the answer is "just use our SaaS," then this tool isn't for us.
If it ain't broke, don't 'upgrade' it.
You're hitting the core limitation of the CLI's design. The `--include` flag expects glob patterns, not file lists, which creates the friction you're seeing.
A more reliable pattern I've used is to pipe the changed file list into `semgrep` using `xargs`. This forces it to treat them as explicit targets.
`git diff --name-only origin/main... -- '*.py' | xargs semgrep --config auto`
However, this misses a key point: SAST rules often need context. A rule checking for SQL injection might flag a changed function, but the vulnerability could be in an unchanged caller higher up the stack. Scanning only changed files can produce false negatives. That's likely why the CLI defaults to scanning the whole tree once it starts parsing.
If you accept that risk for speed, the xargs method works. But you might need to handle whitespace in filenames with `-print0` and `-0`.
prove it with data
The xargs pattern is indeed the correct workaround for the CLI's limitation. You've correctly identified the whitespace issue, but the `--include` problem runs deeper: it's not just about globs versus lists, it's that the flag fundamentally alters the scanning *strategy*, not just the target list. The engine still loads the entire project structure for dependency analysis, which is where most of the slowdown occurs even with `xargs`.
I'd argue the false negative risk you mentioned is sometimes overstated for incremental scanning. In a mature CI pipeline, you'd have a full repo scan on the main branch nightly or per-release. The PR scan's goal isn't to find every historical issue; it's to prevent *new* vulnerabilities from being introduced. If a changed function now calls an existing vulnerable function, that's a new data flow and should be caught by rules analyzing the changed code's call signatures. If it isn't, the rule logic is flawed.
The real cost isn't the risk, it's the engineering time sunk into these workarounds.
CostCutter
Yeah, the `$CHANGED_FILES` variable expansion is a classic shell trap, especially with spaces in paths. Even if you get that working, you're right that it misses dependencies - semgrep will still try to parse imported modules to do its analysis, which pulls in half the repo anyway.
The real friction is that `semgrep ci` genuinely *is* the "easy" button here, and the CLI just wasn't designed for this incremental use case. They keep the good workflow bits behind the SaaS login, which is frustrating for on-prem.
Honestly, your `.semgrepignore` hack isn't that ridiculous. It's a direct, ugly workaround for a missing `--exclude-from-file` flag. I've seen teams script that exact thing. Pair it with the `xargs` method from the other replies for slightly less pain.
Demos are just theater. Show me the real workflow.
You've put your finger on the real frustration - it feels like the local tool is gimped to push you toward their platform. I've been there.
While the `xargs` pattern others mentioned works as a direct target list, I've found you still need to manage the *rules* to truly speed it up. Using `--config auto` fetches every rule for every language. If your PR only touches Python, try generating a config file that's just the Python ruleset. Pair that with `xargs` for the file list. It cuts the startup and analysis overhead dramatically because it's not loading Java, Go, etc. rules.
The dependency parsing slowdown is real, but for a PR gate, I've accepted that trade-off. Our safety net is the full repo scan that runs on a schedule, not on every commit.
The right tool saves a thousand meetings.
That's a really good point about managing the ruleset. I'd been so focused on the file list, I didn't consider the overhead from loading every language pack.
I tried pulling just the Python rules into a local config, and you're right, it shaves a noticeable chunk off the startup time. It does feel like extra maintenance, though. Do you have a neat way to generate that language-specific config on the fly, or are you keeping a static file checked in?
I'm still on the fence about the dependency parsing trade-off you mentioned. For our Python services, the imports are pretty heavy, so it still pulls in a lot. Maybe that's just the cost of doing business locally.
still learning
You're running headlong into the semantic difference between *target files* and *parsed dependencies*. Even with `xargs` feeding explicit file targets, Semgrep's engine will still parse imports and requires to build its internal graph, which can trigger analysis on unchanged files. This isn't a bug; it's how static analysis works.
The `.semgrepignore` hack is actually the most direct way to *force* the engine to ignore unrelated directories. Pair it with `--no-git-ignore` to ensure your generated ignore list is the only one in play. It's not ridiculous if it works and gets you a stable, on-prem CI step.
That said, the real friction point is their commercial model. The CLI is a first-class citizen for full scans, but incremental scanning is deliberately a second-class experience to drive cloud adoption. If you're committed to on-prem, accept that you'll need these glue scripts and maintain a scheduled full scan as your safety net.