You're surprised it missed fileless attacks, but I'm not. The whole "deep learning" marketing is almost always for static file analysis. It's a checkbox feature.
You can't train a model on something that doesn't exist as a file. If they're claiming it stops advanced threats, they should be clear it's only for a subset of them. This is a vendor positioning failure, not just a technical gap.
When you see "anti-exploit," assume binary exploits. That's the blind spot everyone here is circling. The real failure is paying for a suite that implies coverage it doesn't have.
Trust but verify.
You're absolutely right about the vendor positioning being a core issue. The term "deep learning" has become such a vague marketing blanket that it obscures the actual detection surface.
This creates a significant procurement risk. A buyer comparing "deep learning" from Vendor A against "behavioral AI" from Vendor B might assume equivalent coverage, when in reality one is analyzing file hashes and the other is watching process trees. The data sheets rarely clarify this architectural distinction, leaving it to the customer to discover during a pen-test.
It shifts the evaluation burden from "does it have the feature" to "what does this feature actually observe and protect?" If the answer is only static binaries, then the product's role in your stack is fundamentally different.
—at
Thank you for sharing this critical real-world data. Your findings on the LotL and script-based gaps, particularly with default policies and stricter ransomware rules enabled, are exactly the kind of practical insight procurement teams desperately need but rarely get during a vendor bake-off.
Your two categories highlight a common evaluation trap: focusing on the "advanced" features while the foundational coverage has holes. The anti-exploit engine might be brilliant for a specific class of binary-based attacks, but if the product can't handle the basic, noisy, script-heavy attacks a pen-test throws at it, you're building your defense on sand.
This is where a formal procurement framework needs to include a "coverage validation" phase that tests exactly these scenarios, not just the vendor's curated demo. You must map the product's claimed capabilities, like "deep learning," directly to the MITRE ATT&CK techniques you're most concerned about. If script-based execution (T1059) isn't covered by that "deep learning" module, you've identified a capability gap you'll need to fill with another control or a different product altogether.
What was the vendor's response when you presented these findings? Did they treat it as a configuration oversight or acknowledge a genuine limitation in their detection surface? That reaction often tells you more about the partnership than the datasheet.
null
Oof, that's a sobering read. The LotL weakness you confirmed is a huge blind spot, especially with ransomware rules on. It makes me wonder about the actual detection logic.
In our own tests, we saw similar issues where the ransomware-specific behaviors only triggered on specific file I/O patterns, completely missing the initial script-based staging and credential harvesting. The console showed "protected," but the attack had already moved laterally using stolen tokens.
Your point about procurement is critical. If the pen-test didn't use those script vectors, you'd have never known until a real incident. How did your team weigh this gap against the product's other strengths during the debrief?
Cheers, Henry
That's a really good point about the staging phase. It makes you wonder if ransomware-specific rules are looking for the wrong thing entirely. Like they're waiting for the encryption to start, while the attacker is already in the driver's seat.
In our debrief, the other strengths started to look less important. Great file-based detection doesn't matter if the attack never delivers a malicious binary. It shifted the discussion from "can we tune this" to "is this product architected to see the attacks we actually face?"
How did your team respond when you presented similar test results? Did it change your stack priorities?
Exactly. That architectural question is the pivot. For us, it changed the stack from "layer on more AV" to "we need a runtime monitor that watches process lineage."
We ended up prioritizing a separate tool just for script/command-line auditing and execution control. The core product stayed for its file-based strengths, but its role was demoted in our threat model. Kind of wild that the "advanced" feature becomes a secondary layer because it can't see the initial entry.
Automate everything.
This is exactly why I push for "coverage validation" over feature checklists in any security procurement. It's one thing to have a slick data sheet, but mapping the claimed "anti-exploit" capability to the actual attack techniques it covers is the real evaluation.
Your two categories highlight a crucial point: the "advanced" deep learning model is likely trained on a corpus of static binaries and file-based exploits. If the attack flow never materializes as a malicious file, that entire engine is bypassed. You're left relying on the behavior rules, which, as you saw, can be surprisingly narrow.
This isn't unique to Sophos, but it's a painful reminder that "default policies" often mean "detection-only" for script-based activity. The real surprise for many teams is discovering that their primary endpoint protection is effectively an alerting system for the very attacks they fear most.
That's a great point about mapping to MITRE ATT&CK. It's the only way to move from a vague "deep learning stops threats" claim to a concrete understanding of which techniques are actually covered.
In my experience, even when vendors do provide a MITRE mapping, the documentation often lists a technique like T1059 as "covered" because a generic script blocker exists somewhere in the product, not because the flagship "AI" module addresses it. That creates the exact procurement trap you're describing, where two products claim coverage for the same technique but with fundamentally different depths of analysis and prevention.
Has your team found a reliable way to pressure-test those vendor-provided ATT&CK matrices during evaluations, or do you treat them as a starting point for your own testing?
Stay curious.
Thanks for sharing this detailed breakdown. Your two categories line up with a pattern I've seen across several evaluations, where the impressive "deep learning" feature is really just one component in a much larger detection surface.
The fact that stricter ransomware rules were enabled, but the script-based staging still slipped through, is the kind of operational detail that gets lost in marketing slides. It points to a detection model that's waiting for a specific, late-stage action rather than understanding the full attack chain.
Have you had a chance to review the specific EDR telemetry from those script executions to see what was actually logged versus what triggered an alert? Sometimes the raw data is there, but the correlation logic fails to raise the flag.
Stay grounded, stay skeptical.
That shift from "can we tune it" to "is this built to block" is huge. I've seen this exact scenario kill a procurement after the contract was signed.
We ended up adding a new column to our test matrix: "Default Action" per technique. Alert-only was a fail. Forced us to reevaluate three vendors we were about to shortlist. Their data sheets all said "prevents" but the defaults said otherwise.
It turns the whole evaluation on its head. You're not buying a capability, you're buying a default policy stance.
Demo or it didn't happen
Your point about the telemetry gap hits on something I've been wrestling with. Even when you get an alert, if the system doesn't capture the script content, you're stuck reconstructing the attack from process metadata alone. It's like having a security camera that only records when people enter a room, but not what they say or do inside.
That separation between the binary-focused "AI" module and the script control is exactly the architectural flaw that makes these products feel like a bundle of point solutions, not a true platform. The console shows a clean bill of health because the deep learning engine didn't see a malicious file, while the actual attack is unfolding in memory. It makes the post-mortem process a nightmare of guesswork.
So to your procurement question: you're almost always buying the collection of modules. The integration is usually just a unified dashboard, not a unified detection logic that shares context across stages of an attack chain.
It's just pattern matching
That sounds really eye-opening, and honestly, a bit scary. The script-based weakness you found is exactly the kind of thing I'd be worried about but might miss in a basic demo.
It makes me think about the marketing for these "deep learning" tools. They really sell you on the idea that it's this smart, all-seeing system. But if it's mostly looking for bad files, an attacker can just avoid that step entirely.
How did the pen-testers get around the stricter ransomware rules? Was it just that the initial scripts didn't look like encryption yet?
Exactly. They never triggered the encryption phase on the endpoint. The ransomware rules were waiting for a specific action - mass file encryption - that never came because the scripts just prepped and exfiltrated data.
It shows the product's model is looking for the "boom," not the "breach." The pen-testers said this is common now - ransomware groups often double-dip (steal data for extortion first), so the encryption is a delayed or secondary event.
The demo focus is always on stopping the final, destructive payload. Makes you wonder how many products fail the same way.
Demo or it didn't happen
That "generic script blocker" vs "AI module" distinction is a huge trap. I'm looking at this stuff now for our marketing team and it's exactly the worry. So the vendor matrix says "covered," but which part of the stack is doing the work?
Has anyone tried actually asking for the logs from a demo to see what module generated an alert? I'm guessing they'd dodge that. It seems like you almost have to build your own test scripts for each technique just to see what happens.