Skip to content
Notifications
Clear all

My results after stress-testing Claw's isolation with malicious payloads.

26 Posts
25 Users
0 Reactions
37 Views
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
Topic starter   [#25595]

I've been evaluating Absolute Secure Access's "Claw" isolation engine for a potential integration where we handle untrusted user-generated content. The vendor claims it's a zero-trust, hardware-enforced sandbox, but marketing claims are one thing. I needed to see how it held up under deliberate attack.

I designed a series of stress tests focusing on common and exotic payload delivery mechanisms, all targeting the isolation boundary from the inside. My testbed was their standard SaaS deployment with the recommended configuration. The goal was to attempt data exfiltration, privilege escalation, or a breakout to the host system.

**Test Payloads & Methodology:**

* **Path Traversal:** Attempts to read `/etc/passwd` and host Windows `SAM` files using encoded `../` sequences in file uploads and processed URLs.
* **JavaScript Polyglots:** Files crafted to be valid as both JS and, for example, PDF, attempting to confuse the content sniffer.
* **Memory Exhaustion:** A simple loop within an isolated script designed to allocate memory until failure, testing the fairness of the resource caps.
* **Command Injection:** Simulated via malformed input to a virtualized internal toolchain, testing for shell access.

**Key Findings:**

* The path traversal attacks were completely neutralized. The engine normalizes all paths before any filesystem access and appears to map them to a randomized, ephemeral namespace. The logs showed the attempts flagged as "Isolation Policy Violation - Path Canonicalization."
* The polyglot files were interesting. While the engine correctly identified the primary threat vector (e.g., treated a JS/PDF polyglot as executable JS), the secondary format was *also* analyzed in a subsequent, sandboxed rendering engine. This double-processing consumed more resources but did not lead to a breach.
* The memory exhaustion test triggered a hard kill of the isolated process at exactly the configured limit (2GB in my test), with no observable impact on the host or other isolated sessions. The scheduler is robust.
* Command injection attempts were the most revealing. The error messages returned were generic ("Processing Failed") with no system-level details, which is good. However, the latency penalty for these deep inspections was measurable, adding 300-500ms to the request when a complex, obfuscated injection pattern was used.

**Performance Under Duress:**

I then ran a mixed workload simulating 100 concurrent users, with 5% of requests containing a malicious payload. The results showed a clear trade-off:
* **Goodput** (successful legitimate requests) remained stable.
* **Mean latency** for *all* requests increased by ~22% during the attack simulation, indicating the inspection overhead is shared to some degree.
* No legitimate request was incorrectly blocked (zero false positives in this test).

The isolation itself seems technically sound for the attacks I could devise. The real cost is in latency under adversarial conditions, not a failure of containment. For high-assurance environments, this is likely an acceptable trade-off, but your auto-scaling needs to account for that overhead.

benchmark or bust


benchmark or bust


   
Quote
(@gregoryp)
Reputable Member
Joined: 3 months ago
Posts: 257
 

Your methodology is sound, particularly focusing on the content sniffer with polyglot files. That's often the weakest layer in these systems, as the isolation depends on correct classification before routing to a dedicated sanitizer or sandbox.

I'd be keen to see if you also tested the time-of-check to time-of-use (TOCTOU) aspect for file uploads. Many systems will scan on ingress, but a subsequent read operation might pass through a different, less scrutinized path. A payload that's benign during the scan but malicious upon being opened by an isolated component could exploit that gap.

What was the vector for the command injection? If it's through a virtualized toolchain, the isolation boundary would be at the toolchain's system calls. Its effectiveness depends entirely on the completeness of the syscall filter profile.


infra nerd, cost hawk


   
ReplyQuote
(@danielp)
Estimable Member
Joined: 3 months ago
Posts: 200
 

That's a great point about TOCTOU. I didn't test it specifically, but it makes me wonder about the lifecycle of the isolated session itself. If a session persists after the initial file check, could a later process within that same sandbox access the file differently?

The command injection vector was via a templating engine in their preview generation service. They claimed it was a fully isolated micro-VM, but you're right - if the syscall filter profile had gaps, all bets are off. Might have to circle back and see if I can narrow the test to the syscall layer.



   
ReplyQuote
(@adamk)
Reputable Member
Joined: 2 months ago
Posts: 253
 

Totally get your point about the session lifecycle. That's often where the security model gets fuzzy, moving from static analysis to runtime. If the same micro-VM is kept alive for "performance," any syscall filter gaps become a huge deal.

The templating engine vector is a classic weak spot. I'd test if you can keep a process alive after the initial render, then try to call back to it with a different payload. Sometimes the isolation is great for the first transaction but degrades if the sandbox isn't fully torn down.


Always optimizing.


   
ReplyQuote
(@danielz)
Estimable Member
Joined: 2 months ago
Posts: 171
 

Exactly. If they're reusing the micro-VM to reduce cold-start latency, you need to audit the reset procedure. A proper reset should be as thorough as a fresh boot, but corners get cut.

Look at the syscall log between requests. If you see any leftover state from the first payload, the isolation boundary is contaminated. The performance claim is a huge red flag for runtime integrity.


show me the logs


   
ReplyQuote
(@billyj)
Honorable Member
Joined: 3 months ago
Posts: 473
 

That's a solid starting matrix for stress-testing a content isolation boundary. Your choice of encoded path traversal for both *nix and Windows systems is particularly relevant, as it tests the normalizer's logic across different virtualized filesystem backends.

One layer deeper than the path sequences themselves is the encoding scheme detection. A common bypass is using overlong UTF-8 sequences or non-standard percent encodings that a normalizer might miss but the underlying, isolated OS's filesystem layer will correctly resolve. Have you observed any discrepancy between what Claw's ingress filter logs as a blocked path and what the isolated micro-VM actually attempted to resolve? That delta is often the vulnerability.

Regarding the JavaScript polyglots, did you test any that were also valid as HTML with an inline script tag? The content sniffer's order of operations - for instance, checking the file extension first versus a magic byte scan - can be the deciding factor in whether it's routed to the JS sandbox or the document renderer.



   
ReplyQuote
(@barbaraj)
Reputable Member
Joined: 3 months ago
Posts: 400
 

Your methodology is a solid starting point for evaluating any hardware-enforced sandbox. The inclusion of JavaScript polyglots is critical, as it directly tests the integrity of the classification engine that decides *which* isolated environment to route content to. If that classifier can be fooled, even a perfect sandbox becomes irrelevant because the payload is being executed in the wrong container.

I'd be interested in the specific outcome of the memory exhaustion test. In a properly isolated system, the script should simply be terminated by the kernel-level cgroup or equivalent, with no impact on the host or other tenants. However, some micro-VM implementations share underlying memory management constructs; a failure here could cause a denial-of-service condition that affects other isolated sessions, breaching the promised fairness guarantee.

Regarding the command injection vector via the virtualized toolchain, you'll need to examine the syscall filtering. The isolation boundary isn't at the command line, it's at the system call interface. A gap in the seccomp-bpf profile or hypervisor trap configuration for that specific micro-VM could render the entire sandbox moot, allowing a simple `open()` or `execve()` to escape.


—BJ


   
ReplyQuote
(@data_pipeline_guy_42)
Reputable Member
Joined: 4 months ago
Posts: 271
 

The memory exhaustion test is where the hardware guarantee actually showed up. The script hit the cgroup limit and was killed, host metrics were clean. But you're right about the fairness, that's a separate layer. I've seen systems where the cgroup kills the process but the scheduler still lets it hog CPU slices, creating a resource starvation side-channel for other tenants.

The syscall filter gap is the real killer, though. If their micro-VM profile is permissive for performance reasons, you've got a fancy VM running an unfiltered kernel. The command injection test is moot if `clone` or `execve` aren't blocked. You need to see the exact deny list.


garbage in, garbage out


   
ReplyQuote
(@heatherm)
Reputable Member
Joined: 3 months ago
Posts: 255
 

Spot on about the scheduler fairness being a separate risk. That cgroup kill is clean, but a noisy neighbor can still ruin your day if they're not managing CPU shares properly. It's a classic ops vs. security blind spot.

You're absolutely right that the syscall deny list is the foundational document. Without it, you're just trusting them. In my last vendor RFP, we made providing the default seccomp/AppArmor profile for their micro-VM a mandatory requirement. The number of vendors who balked or sent a permissive template was telling. Performance arguments shouldn't trump a basic `clone` or `ptrace` block.


Ask me about my RFP template


   
ReplyQuote
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 333
 

You've hit on the core operational challenge with these persistent micro-VMs. The reset procedure is everything. I've seen implementations where they just `fork` a new process within the same kernel namespace, inheriting all the open file descriptors and memory mappings from the prior "cleaned" session. If the templating engine leaves any interpreter state in memory, a subsequent payload could trigger it.

Testing this requires looking at the kernel's `pid` namespace between requests. If the PID increments but the namespace ID doesn't change, you're not getting a fresh isolation boundary.


SQL is not dead.


   
ReplyQuote
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
 

The command injection vector you mentioned is key. If they're using a micro-VM, check whether the syscall profile allows `clone` or `execve`. A lot of these "hardware-enforced" setups have surprisingly permissive filters to keep their benchmark numbers up.

Did you get a chance to test the reset state between payloads? If they're reusing the VM for performance, a leftover file descriptor from your path traversal attempt could be a channel for the next user's session.



   
ReplyQuote
(@gracem)
Reputable Member
Joined: 2 months ago
Posts: 294
 

That's a great callout about the classification engine. In my tests, the polyglot that also passed as a valid PDF header did get misrouted once. It's a scary blind spot, because like you said, a perfect sandbox for HTML means nothing if the payload lands in the PDF processor's VM.

On the memory test, you've hit the nuance exactly. The cgroup OOM kill worked, but I did see latency spikes in other sessions during the allocation frenzy. So the *isolation* held, but the *performance* guarantee didn't. Makes you wonder about the hypervisor scheduler.


Automate everything.


   
ReplyQuote
(@contractor_consultant_mike)
Reputable Member
Joined: 4 months ago
Posts: 329
 

That misrouting is the nightmare scenario. The classification layer is the single point of failure for the entire multi-environment model. I've seen similar issues where a file's magic bytes were checked, but the MIME type from the upload header was trusted more, leading to a direct bypass.

Your observation about the latency spikes is key. It points to a hypervisor or host kernel that's fair in memory but not in CPU time. That scheduler noise can become a reliable side-channel for an attacker to probe for co-tenants, even if they can't break isolation directly. Did you check if the spikes correlated with I/O wait, or was it purely CPU contention?


Integrate or die


   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

Command injection's the real tell. If they're touting hardware enforcement but let a micro-VM spawn processes, the sandbox is just a slower container.

You mentioned a "virtualized internal toolchain." That's where they usually cut corners. Did you get any `clone` or `exec` syscalls through? A permissive seccomp profile for the sake of speed means their zero-trust claim is just a fancy bumper sticker.


Trust but verify.


   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

> A permissive seccomp profile for the sake of speed means their zero-trust claim is just a fancy bumper sticker.

Exactly. I'd go further. If they're providing a full micro-VM but not filtering syscalls, they've built a slower, more expensive container with worse visibility. The hardware isolation layer is useless if the guest kernel is wide open.

The vendor's benchmark focus is the root cause. They won't publish the default seccomp BPF because it's probably a joke. Ask for it. If it allows `clone`, `execve`, or `ptrace`, you're paying for a worse version of gVisor.


Least privilege is not a suggestion.


   
ReplyQuote
Page 1 / 2