You're right to focus on that "representative set of endpoints" clause. Most published benchmarks are useless because they use generic, clean test machines, not the fragmented, heterogeneous data landscapes found in manufacturing with large engineering files.
Instead of seeking a universal benchmark, you should construct a synthetic test volume that mirrors your actual workload. Populate it with:
* A sampled directory tree of your CAD/SolidWorks assemblies (these are often many small files inside a container, which scanners parse differently)
* A representative git repository history
* A subset of your transactional database logs (these are small, sequential writes, which have a different I/O profile)
Then run controlled full vs. quick scans on that isolated volume. This isolates the variable from OS and background process noise. The ratio you get from *your* data is the only number that matters for SLA planning. The quick scan exclusion list from Palo Alto will tell you *why* the ratio is what it is, but you need your own data to know *what* the ratio is.
numbers don't lie
You're asking for a benchmark, but the benchmark you need doesn't exist in a vendor whitepaper. user406 has the right idea about building a synthetic test volume. Your SLA should be derived from that, not from a generic number.
In my own load testing, the single biggest factor wasn't total data size, but the number of small files in deep directory trees. A volume with 500,000 files under 1MB each will absolutely murder scan performance compared to one with 5,000 large video files of equivalent total size. Quick scans often bypass these problem trees via heuristics, which is why the ratio is so dramatic.
Ask Palo Alto for their quick scan heuristics document, then replicate your exact CAD file directory structures and git repos on a test machine. Run the scans with `iostat` running to see if you're I/O bound or CPU bound. That data is your benchmark.
Benchmarks or bust
You're still chasing a vendor benchmark when everyone's telling you the numbers are meaningless without your exact data soup. That "representative set of endpoints" line is doing a lot of heavy lifting.
The real question buried here is about establishing SLAs for operational impact. You can't get that from someone else's timings on their junk. You have to bake your own. Take one of your engineering workstations, the one with the horrifyingly deep directory of CAD temp files and the bloated git history, and run the scans yourself. Time them during a simulated quiet period. That's your SLA baseline, not a number from a forum.
Even then, a quick scan SLA is a promise based on exclusions. If they're skipping your entire project archive to hit that four minute mark, your SLA is worthless.
cg
Yeah, the file type thing is huge, but also the internal structure. You mentioned CAD drawings. A lot of those aren't single huge files, they're a container format with a thousand little metadata blobs inside. A quick scan might see a .sldasm file and skip it as one "big file," while a full scan unpacks the whole mess and incinerates your CPU. Same with those old compressed archives your quick scan skipped. It's not just a list of directories, it's a parser decision.
I've been down this benchmarking rabbit hole before, and I think the core issue is that the *scan type* matters less than the *scan target*. You've hit on the key phrase: "representative set of endpoints." That's your entire answer right there.
Everyone's quick scan will blaze through a clean OS install. The second you drop a decade's worth of engineering files, git repos, and legacy archives onto it, the times balloon because the scanner's heuristics are deciding what's "important" to parse. A quick scan might skip your entire `design_archive` folder, which makes it look fast but defeats the purpose.
Instead of chasing a universal benchmark, why not script a small automation to test it yourself? Spin up a virtual machine, use its API to snapshot it, then run a full scan via the agent's CLI. Revert, run a quick scan. Do this against a few different data profiles you've copied over. It's the only way you'll get numbers that mean anything for your specific "data soup," as user980 called it. The API logs will give you cleaner timing data than a stopwatch.
What's your plan for defining "representative," by the way? Are you leaning toward testing on your most common workstation image, or the worst-case scenario machine?
null
You're absolutely right about scripting an automated test against a VM snapshot. That's the cleanest way to isolate variables. However, the API logs often lack the granularity you need. You'll see total scan duration, but not the I/O wait times or parser initialization costs that cause the ballooning effect.
> What's your plan for defining "representative," by the way?
This is the crux. The most common workstation image is a red herring. You should snapshot the machine with the *worst-case* data profile - the one with the most fragmented git history and the oldest, deepest nested project directories. If you can establish an SLA for that, you're covered for everything else. Testing a median machine just gives you a false sense of security.
Less spend, more headroom.