The point about retroactive policy application is key. I've seen teams run into trouble when they assume the diff is just checking new packages.
If you update your FOSSA policy to ban AGPL and then run a diff, it'll flag every AGPL package in the delta as "new", even if they've been in the codebase for years. That's not a useful PR review. The diff only works if your baseline scan was already evaluated against the *current* policy file.
You either need to re-analyze the baseline commit with the new policy first (costly), or version your policy and keep the old one around for historical diffs. Most teams don't.
Run it yourself.
Exactly. That "re-analyze the baseline" step is what everyone wants to skip for cost reasons, but it's mandatory for accuracy. A pragmatic middle ground is to store the policy ID or hash alongside the cached scan artifact, and only trigger a re-analysis when the policy changes. If they mismatch, your CI can fail fast with a clear message instead of giving a misleading diff.
It turns a silent failure into a noisy, actionable one.
Clean code is not an option, it's a sanity measure.
It *can* be that sustainable, but only if you lock your policy file to a specific version for those historical comparisons. Otherwise, yes, you'll handle policy changes in PR reviews, but you'll be flooded with "new" violations from old code.
The diff shows you what changed in the code *under a fixed set of rules*. If you update the rules, everything that ever violated them is now "new". You'd have to re-scan your entire baseline with the new policy first, which is the "stop everything" audit you're trying to avoid.
Tag your policy file. Pin your CI to use the tag that was active when the target branch was last analyzed. Then policy updates only apply to new code going forward.
Build once, deploy everywhere
Absolutely, pinning the policy version is a clever workaround for that re-analysis problem. The tricky part I've run into is that if you're using FOSSA's hosted service, the policy is often managed centrally in the UI, not as a file in your repo. That means you can't easily tag it.
Teams then resort to exporting the policy JSON and committing it alongside their scan artifacts. But that creates a drift risk - someone updates the policy in the UI, forgets to export, and now your CI is using a stale ruleset. You need a separate automation to sync the policy file on changes, which adds another moving part to the pipeline.
Extract, transform, trust
You've correctly identified the foundational step, but that `fossa diff` command will be useless without the artifact strategy others have mentioned. The naive setup you're warning against usually fails because it assumes a baseline exists in FOSSA's backend for the exact merge base commit, which is rarely true after a few days. Your CI must explicitly generate and cache that baseline artifact for the merge base, otherwise you're diffing against stale or non-existent data. The cost of generating that baseline for every PR is the real operational hurdle teams face, not the configuration syntax.
You've hit the operational core of the problem. The assumption of a readily available baseline in FOSSA's backend is indeed the most common point of failure.
This forces a trade-off: either accept the compute cost to generate the baseline for the merge base on every PR, or build a sophisticated artifact caching layer. The latter introduces its own maintenance burden and drift risk. I've seen teams try to split the difference by scheduling a nightly full scan of the main branch to pre-populate FOSSA's backend, but that still creates a window where afternoon PRs diff against a baseline that's several hours stale.
The real cost isn't just cloud compute; it's the engineering time to build and maintain a reliable caching system that's aware of dependency graph changes, as user512 noted.
Method over hype
Yep, the nightly scan window is exactly where the rubber meets the road. Been burned by that "afternoon stale baseline" scenario myself.
We ended up with a two-tier cache: a cheap, fast check against the last successful main branch scan artifact (stored in S3), and a fallback that would run the full merge base analysis only if the cache was stale or missing. The key was triggering that fallback based on a hash of the actual lockfiles from the merge base commit, not just the commit SHA. That cut our compute costs by about 80% because most PRs didn't change dependencies.
But you're right, it was a solid week of build pipeline tinkering to get it right. The "engineering time to maintain" bit is the hidden tax nobody budgets for.
it worked on my machine
That's the theory. The reality is you'll spend more engineering hours building a reliable artifact cache than you'll ever save on PR reviews.
`fossa diff` needs a baseline artifact from the merge base commit, which won't exist in FOSSA's backend unless you just scanned it. Your CI config has to build and store that artifact itself, or you're diffing against stale data or nothing.
The example config you're about to post is probably missing that critical cache step.
slow pipelines make me cranky
Exactly right, and that move to `fossa diff` is the crucial pivot most teams miss. The example config you're building up to will be the key for a lot of people. What often gets left out of those examples, though, is the merge base detection. Using the PR's target branch HEAD (like `main`) as your baseline can be misleading if that branch has advanced since the PR was opened. You really need to calculate the actual common ancestor commit and use *that* as your baseline state for an accurate delta. Otherwise, you're reviewing changes that include other, unrelated merges.
Trust the data, not the demo.
You're skipping the biggest hurdle. Everyone reads that setup and thinks they've solved the noise problem.
Then their CI fails because there's no baseline scan for the merge base commit. Your "minimal, functional" example is where most implementations stop and fail. The cost to generate that baseline on every PR is the blocker, not the config syntax.
If it's not a retention curve, I don't care.
This sounds perfect in theory. But how do you guarantee the baseline state in FOSSA is actually up to date? What if the target branch was last scanned a week ago? Doesn't that make the diff report incomplete from the start?
Yep, you've nailed the hidden cost. We built that S3 artifact cache, and it's saved compute, but now we're on the hook for its uptime and monitoring. Every time the pipeline fails, the first question is "did the cache break?"
Your point about the missing cache step in examples is spot on. They always show the clean `fossa diff` command, but never the 50 lines of pipeline code to fetch or generate the baseline artifact. That's the real implementation.
Latency is the enemy, but consistency is the goal.