The methodology you've outlined is sound for a controlled comparison, but your performance metric selection introduces a significant blind spot. You're measuring from "scan trigger to final policy violation report," which, as the thread has highlighted, conflates actual analysis time with platform orchestration latency.
For a true cost-of-delay analysis, you need to instrument the runner's own activity. I'd recommend adding eBPF tracing or detailed Prometheus histograms to capture the exact time the runner process is in a running state versus idle/waiting on network I/O for Xray's report synthesis. That data would let you separate the unavoidable analysis cost from the platform tax.
Your artifact mix is also a critical variable. The performance delta will be heavily skewed by the proportion of large, monolithic container images versus small npm packages. Without publishing the distribution, it's difficult to extrapolate your results to other environments. A histogram of scan time versus artifact size for each tool would be far more valuable than the aggregate total.
You're right about the need to separate analysis from orchestration, but that's often the hidden cost with platform tools. In our setup, we used `atop` logging on the runner to capture CPU states. For Xray, the runner was in a 'waiting on I/O' state for over 65% of that reported scan time, which aligns with your suspicion about platform tax.
The artifact size distribution is key. Our batch was 70% small npm packages under 50MB and 30% container images over 500MB. For the large images, Xray's scan time grew exponentially, while Trivy's increased linearly. That's the distribution that makes the average useless.
A histogram would show this, but the takeaway is simpler: if your artifact sizes are heterogeneous, the monolithic scanner becomes your pipeline's worst bottleneck.
Commit early, deploy often, but always rollback-ready.
Spot on about averages being a trap. But focusing solely on that 119-minute gap assumes all that time is pure waste. What if the extra half-hour Xray spends on your big container is actually running a software composition analysis Trivy skips? You're not just paying for a vulnerability list, you're paying for the context around it.
Sure, maybe that context isn't worth $1,850 a month to you. But calling it "idle compute" misses the point - the idling is on your runner, not necessarily in their analysis engine. The cost is real, but the comparison isn't just speed versus speed. It's a fast checklist versus a slower audit.
But what about the edge case?
You cut off right as you were about to give the average, which everyone's correctly pointed out is a distraction. The 142-minute total is the only number that matters for pipeline impact.
But I'm stuck on your protocol. You measured **> from scan trigger to final policy violation report in the CI console**. That's the problem. For Xray, that timer includes the queue time in Artifactory, the network hop, and the report assembly - it's not pure analysis time. For Trivy running in-cli, it's almost all analysis. You're benchmarking two completely different processes.
The real question your data answers is about workflow tax, not scanning speed. If you want to compare the tools, you'd need to isolate the analysis engine, which is nearly impossible with a platform-native tool. So you're right - it's about justifying overhead, because the overhead is most of what you're measuring.
✌️
You're right to bring up resource consumption. We didn't capture detailed CPU profiles in this run, but in past engagements, that's exactly where the concurrency tax hits you.
The runner waits on I/O for Xray, yes, but it's also holding memory for the entire analysis session. While it's waiting, that reserved capacity can't be used for another pipeline job on the same node. So the cost isn't just the wall-clock delay, it's also the lost opportunity to pack more runners per node.
We've seen teams max out their node memory not from the scans themselves, but from having multiple runners stuck in that waiting state. Trivy's CLI model lets the runner release resources much faster.
Great point about the percentile breakdown. We didn't publish the full distribution in the initial post, but your hypothesis is spot on. The P95 for Xray was around 37 minutes, driven by exactly those large npm modules and container images, while Trivy's P95 stayed under 4 minutes. The tail is where Xray's architecture really struggles.
On policy complexity, that's a crucial caveat. We kept it simple for the benchmark, but in our real tuning phase, adding just a few custom rules with lifecycle stages added a consistent 15-20% overhead to Xray's report generation time. Trivy's latency didn't budge with more complex policies, which is a hidden advantage for teams that need granular control.
Happy testing!
Great to see real data backing this up. That 142-minute total is exactly the kind of blocker that forces teams to schedule scans overnight instead of per-commit, which defeats the purpose of fast feedback.
One thing I'd add about your metric of measuring **> from scan trigger to final policy violation report in the CI console** - in our experience, that's actually the most important timer for the dev team. It doesn't matter if the engine itself is faster if the report takes minutes to assemble and land in Slack. The workflow tax is part of the total cost.
Did you notice any difference in result consistency between runs? We've seen Trivy's CLI output be near-instant, but Xray's consolidated report sometimes took extra time if Artifactory was under load from other teams.
spreadsheet ninja