Yeah, the `.pom` file pattern is dead-on for us too. Your pre-commit check and filtered CI scan sound smart.
Our biggest gap with the type-based exclusion is when a runtime jar pulls in a non-test-scoped parent with a GPL declaration. That still sneaks through. We ended up building a small CLI that dumps the full effective POM for each flagged artifact, which sometimes shows the problematic parent declaration the scanner latched onto.
>It's not perfect, but it cuts the false positives by about 80%
That's about our experience as well. The remaining 20% are the real time sinks, chasing phantom licenses through metadata.
Latency is the enemy, but consistency is the goal.
That CLI approach for dumping the effective POM is a solid move. It's essentially replicating what the scanner should be doing in its resolution phase.
We found the same gap with runtime jars and non-test parent declarations. In our case, we started logging the full Maven dependency tree alongside the scanner output and correlating them. A surprising number of those remaining 20% false positives were artifacts where the parent POM's license declaration was `GPL-2.0-only` but the child's own `pom.xml` explicitly overrode it with something like `Apache-2.0`. The scanner was taking the first license label it found and stopping its crawl.
Have you considered feeding your CLI's effective POM output back into the scanner as a source of truth? It's extra work, but it could bridge that metadata gap.
Ignoring internal repo paths is a good stopgap, but it just masks the problem for releases. We ended up with GPL flags on artifacts pushed to a public repository because the scanner crawled metadata from a mirrored copy of Maven Central. The watch filter doesn't help you there.
You're right about the tool not being smart. It's scanning metadata, not code, and calling that a result.
show me the logs
The pattern you've described with transitive dependencies is a common failure mode in license scanning, but I'd argue the root cause is even more fundamental than classifier or license file parsing. The scanner is operating on a flawed assumption: that dependency metadata constitutes a legally binding declaration, which it often does not.
In my experience, the scanner's dependency graph resolution frequently conflates build-time and runtime dependency trees. A library's POM might list a GPL-licensed tool for code generation under its 'plugins' section, and the scanner incorrectly attributes that license to the library's runtime artifact. This is why you see it on Docker images and Maven artifacts alike. The scanner isn't distinguishing between a component needed to *produce* the artifact and a component that *is* the artifact.
Have you tried correlating the false positives with the presence of specific Maven plugins or build extensions in the dependency tree? That's where we found a significant cluster of our GPL flags.
SQL is not dead.
Forensic examination is a good start, but have you actually pulled the incident report from your last build failure? The audit trail should show the exact metadata path Xray followed.
Your curated library set is irrelevant if the scanner is crawling parent POMs from your mirrored Central. The false positive isn't about your vetting; it's about Xray's flawed dependency resolution conflating build and runtime scopes. I'd bet your "common, permissively-licensed library" has a parent or plugin declaration somewhere up the tree.
>wasting considerable effort
That's the real cost. You're not just debugging a scanner, you're manually auditing because the tool can't distinguish a license file from a pom.xml declaration. Have you tried comparing the effective POM of the flagged artifact against Xray's claimed violation? The mismatch is usually obvious.
- Nina
Exactly. That CLI to dump the effective POM is a great idea. We do something similar by checking the 'mvn help:effective-pom' output for flagged items.
It helped us spot a case where a common logging library had a parent POM with a GPL-2.0-only declaration from like 2012. The scanner grabbed that, even though the library's own license is Apache. It's all in that metadata ancestry.
measure twice, ship once
It sounds like you're doing all the right groundwork with your curated library set, so that persistent GPL flag must be especially frustrating. Your forensic approach of tracing it to transitive dependencies is the key. A lot of folks hit this same wall.
The classifier and license file interpretation you mentioned can definitely be part of the problem, but in our experience, the scanner often stops its analysis too early. It finds a license tag in a parent POM and doesn't continue to see if the immediate dependency's own POM overrides it with something like Apache 2.0. That might be what's happening with your common library.
Have you been able to run a local `mvn dependency:tree` or check an effective POM for one of these flagged images to see the exact path? That usually reveals the metadata culprit.
—HR
The irony of a tool that's supposed to save you time creating this much forensic work is pretty rich. You've done the right prep work with a curated, vetted set of libraries, and yet you're still stuck playing detective in your own private repository.
What's really galling is that you're paying for this, probably a hefty enterprise subscription, and it's making you do the manual audit work it was sold to automate. Tracing classifier misinterpretations feels like debugging the scanner's homework instead of getting results.
Has anyone on your team calculated the cost of this "considerable effort" against Xray's licensing fee? I'd be curious to see if the math still works out in their favor.
—DW
That calculation is a valid, and often sobering, exercise. However, the goal shouldn't be to justify the tool's cost, but to quantify the problem for the vendor. We've had success framing it as a blocker to scaling adoption internally. When teams see the manual audit overhead, they resist new security and compliance initiatives.
Presenting that effort as a concrete inefficiency - hours spent per week chasing false positives - shifts the conversation from "is the tool working?" to "this is what we need fixed to realize the value we're paying for." It puts the onus on support to provide a resolution path beyond workarounds.
—daniel
That's a good point about making it a scaling issue. I tried to tell our security team about the manual hours, but they said it was just part of the compliance process. Framing it as a blocker to wider adoption might get their attention.
How did you present that data? Just a weekly hours tally, or something more specific?