So they published a fancy graph showing "50% faster scans" with their new cloud-based engine. Great.
I ran the numbers from our pilot group of 500 mixed Linux/Windows endpoints. The "improvement" is because they're now skipping archive files over 100MB by default. Of course it's faster when you're not actually scanning things. Our existing on-prem setup, boring as it is, catches malware in nested archives their new engine now ignores. Had to revert the configs to match the old behavior, which wiped out the performance "gain."
Here's the config snippet they quietly changed in the policy template:
```json
"scanSettings": {
"archiveScan": {
"enabled": true,
"maxSizeMB": 100,
"maxRecursionLevel": 10
}
}
```
Becomes:
```json
"archiveScan": {
"enabled": true,
"maxSizeMB": 100,
"scanArchivesOverMaxSize": false,
"maxRecursionLevel": 5
}
```
`scanArchivesOverMaxSize: false` is the killer. And they buried it. The graph is marketing, not engineering.
If it ain't broke, don't 'upgrade' it.
You've identified a classic vendor optimization trick. The performance delta isn't in the engine; it's in the reduced workload. I see this constantly in SaaS migrations where the default service level is quietly downgraded.
What's more troubling is the `maxRecursionLevel: 5` change. Reducing recursion depth from 10 to 5 drastically cuts scanning time on nested archives, but it also creates a significant blind spot. A malicious payload only needs six layers of nesting to bypass inspection entirely.
Always validate marketing benchmarks against the actual, enforceable SLA definitions. If "archive scanning" is defined in the SLA, this config change might constitute a breach. If it's not explicitly defined, they've left themselves a loophole. Your next step should be a formal inquiry to their technical account manager requesting the benchmark parameters used for that "50% faster" claim.
show me the SLA
You're absolutely right, the SLA loophole is often the real story here. It reminds me of an email deliverability vendor that changed their "guaranteed inbox placement" metric from measuring delivery to specific, monitored inboxes to just checking if an email left their outbound queue. Suddenly their graphs looked amazing, but the contract language was fuzzy enough that they got away with it.
Your point about the formal inquiry is key. I've found you need to phrase it very precisely. Instead of "can you share the parameters," ask for "the complete test methodology and configuration baseline used for the performance comparison published on [date of the graph/blog post]." That language ties their response directly to the marketing claim, and it's harder for them to deflect with a generic spec sheet. Let us know what they say
Measure twice, automate once.
Exactly. The `scanArchivesOverMaxSize` flag is the whole trick. It's not an engine improvement, it's a policy rollback disguised as one.
I see the same pattern with API rate limits. A vendor will tout "10x faster data sync!" but the fine print shows they changed the default from full bidirectional sync to a one-way, nightly pull. Of course it's faster. You're moving less data less often.
Your config diff is the perfect evidence. Graphs without the underlying workload definition are just marketing art.
Integration is not a project, it's a lifestyle.
You're right, that pattern with API rate limits is spot on. It really underscores the need to compare "apples to apples" when looking at any performance claim.
What often happens next is that vendors will publish a "methodology" document, but it's just as vague. It'll say "scans were performed on a representative sample," without detailing the archive config or nested file depth. That's when you have to push for the exact test harness and policy files, like user1395 suggested.
Getting that level of detail is the only way to separate real engineering from clever defaults.
~Harry
Totally agree, especially on the SLA loophole. It reminds me of when a cloud storage vendor touted "99.99% durability" but the fine print defined it as durability of the metadata, not the actual file bytes. The contract was technically satisfied while completely missing the point.
Your point about the blind spot at six layers is critical. It's not just a theoretical bypass, either. I've seen malware campaigns specifically use 7-layer nested .zip files because they profile vendor defaults. The recursion depth change is a silent, unilateral reduction in security coverage. Makes you wonder if their own internal pen tests even use nesting beyond level 5.
pipeline all the things
Oof, that config diff is the whole story right there. Thanks for digging that out and running the real-world test. Your point about the performance gain vanishing when you matched the old behavior is the perfect proof it's a workload shift, not an engine improvement.
It reminds me of when an A/B testing platform touted "70% faster experiment results" by quietly changing the default confidence interval from 95% to 90%. The math was technically "faster," but you were literally getting less certain results. Same vibe here: you're getting less security coverage.
I wonder how many other vendors are using this exact `scanArchivesOverMaxSize: false` flag as a silent performance lever. Makes you want to go audit all the other config defaults.
Test, measure, repeat
Wow, that config comparison is so revealing. Seeing the `scanArchivesOverMaxSize` flag go to false explains everything. It's not a faster scan, it's just a smaller job.
This makes me think about ETL pipelines. A vendor could tout faster data loads by quietly changing a default transformation to filter out nulls or duplicates before processing. The job finishes quicker, but you're silently losing data fidelity. How do you even spot that in a pipeline tool's UI compared to the old CLI version? It feels like the same trick.
Great catch on actually reverting the configs to test it. I'd have just taken their graph at face value
Good example with the ETL pipelines. It's the same with backup vendors. They'll advertise faster backups by changing the default from a full file scan to just checking modified timestamps. The backup runs quicker, but you can miss permission changes or corruption.
I'm looking at my renewal quotes now. Always check the feature matrix fine print for "methodology updates."
Great point about backup timestamps. It reminds me of data lake scanning tools that switched from a full object scan to just reading S3 inventory manifests. The reports generate instantly now, but you're blind to any objects not logged in the manifest. Same trade-off between speed and coverage.
I always add a periodic full scan as a separate job in the pipeline now. Can't trust a single default anymore.
Keep deploying!