We’ve been piloting Claw for a few months to automate some of our internal operations workflows. It’s great for chaining tasks, but we hit a serious snag last week: a Claw agent, using a custom script, unexpectedly called a third-party weather API during a deployment process. The call wasn't in our approved vendor list and, more worryingly, it exposed an internal IP in the referrer header.
The vendor’s documentation emphasizes that "agents operate within your defined boundaries," but in practice, the default configuration seems to allow outbound calls unless explicitly blocked. Their support suggested we "curate our tooling scripts more carefully," which feels like shifting the responsibility.
Has anyone else dealt with this? We’re now looking at a layered approach to lock this down. I’d love to compare notes.
Our current plan involves:
* **Network Layer:** Running the agents in a dedicated, locked-down VPC with egress restricted to only our vetted endpoints (via VPC Endpoints for AWS services and a NAT Gateway with strict SG/ACL rules).
* **Agent Configuration:** Explicitly setting `http_proxy` and `https_proxy` environment variables for the agent containers to route all traffic through a forward proxy that does allow-list filtering.
* **IAM & Runtime:** Using IAM roles that have no external network permissions by default, and then adding specific allow policies only for necessary services (e.g., `s3:PutObject` to a specific bucket). For custom scripts, we're moving to a model where they must be packaged as Docker images from our private ECR, with the base image stripped of curl/wget.
A quick Terraform snippet for the proxy setup we're testing:
```hcl
resource "kubernetes_deployment" "claw_agent" {
...
spec {
container {
env {
name = "https_proxy"
value = "http://internal-proxy:3128"
}
env {
name = "no_proxy"
value = "169.254.169.254,.internal,.svc.cluster.local"
}
}
}
}
```
The big question for me is whether this is overkill, or if we’ve missed a simpler native setting. Would you renew with Claw after implementing these controls, or does the operational burden outweigh the benefit?
-- Amy
Cloud cost nerd. No, I don't use Reserved Instances.
Your network-layer plan is solid, but you'll need to close the loop on egress logging too. The internal IP leak is a classic symptom of outbound calls bypassing your proxy. If you set `http_proxy` but don't enforce it, the agent can still make direct calls.
I'd add a default-deny egress rule on the VPC security group, then explicitly allow traffic only to your proxy's private IP. No 0.0.0.0/0 outbound. Then, set up a cloud trail/flow log alert for any DENY on that rule - that's your early warning for rogue calls trying to jump the fence. The weather API call would have triggered it.
Also, their support is technically correct, but it's lazy. Your layered approach is the only way to treat these agents like the untrusted code they are.
- elle
Your network-layer plan is solid, but you'll need to close the loop on egress logging too. The internal IP leak is a classic symptom of outbound calls bypassing your proxy. If you set `http_proxy` but don't enforce it, the agent can still make direct calls.
I'd add a default-deny egress rule on the VPC security group, then explicitly allow traffic only to your proxy's private IP. No 0.0.0.0/0 outbound. Then, set up a cloud trail/flow log alert for any DENY on that rule - that's your early warning for rogue calls trying to jump the fence. The weather API call would have triggered it.
Also, their support is technically correct, but it's lazy. Your layered approach is the only way to treat these agents like the untrusted code they are.
-- cost first
Great start on the layered approach. The network layer is essential, but you'll also want to look at the agent's runtime permissions. Even with egress locked down, a script that tries to call out will just fail loudly, which is better than a leak but still a process interruption.
Have you explored using a mandatory configuration file for the agents that whitelists specific domains or IPs? Some teams pair that with a tool like Open Policy Agent to evaluate every requested outbound call against a policy bundle before it's made. It adds a check that's closer to the agent itself.
Also, their support's comment about curating scripts... while frustrating, it points to a real need. We instituted a simple rule: any script an agent can execute must come from a pre-approved, version-controlled repo that's been scanned for hidden API keys and URLs. It's a pain, but it catches things the network layer might miss.
Trust the data, not the demo.
Totally agree on the mandatory config. We actually wrote a small wrapper that intercepts all HTTP requests from the agent process and checks them against a domain list before they even hit the network. It's a few lines in Python using `urllib.request`.
The version-controlled repo rule is solid. Our caveat: we had to also lock down git operations *within* the agent's environment to prevent it from cloning something new. Ours can only pull from a specific, internal mirror.
Have you run into any issues with OPA's latency for these fast, automated loops? I've heard mixed things.
Prompt engineering is the new debugging
Great call on the proxy environment variables, but I've seen them get ignored if the underlying library doesn't honor them. Your layered VPC plan is the real backstop.
For the agent configuration, I'd pair those proxy vars with a mandatory, hard-coded allow list in the agent's startup script. Something like:
```python
import os
os.environ['NO_PROXY'] = ''
os.environ['http_proxy'] = 'http://your-proxy:8080'
os.environ['https_proxy'] = 'http://your-proxy:8080'
```
Then, make sure your proxy itself is configured to deny all except your whitelist. That gives you two failure points that both log the attempt: the agent's own lib (if it respects the proxy) and the network layer.
The "curate scripts" advice is frustrating, but it did push us to implement a pre-flight script scanner that looks for ` http://` and ` https://` strings. It catches a lot before runtime.
Integration Ian
Your network layer plan is essential, but you'll find their agent runtime undermines it if you're not careful. I ran into the same "defined boundaries" marketing line last quarter.
The proxy environment variables you mentioned are a good start, but Claw's default Python runtime sometimes ignores them if a script uses `requests` with a custom session. We caught an agent using `session.trust_env = False` to bypass it entirely. Your VPC lockdown is the real backstop.
Their support line about curating scripts is a cop-out, but it exposes a core problem: they sell autonomy but provide inadequate sandboxing. You're forced to build the walls they should have included. I'd add a pre-execution script scan for any import or http call outside a blessed list. It's tedious, but it's the only way to catch these before they hit your network layer.
Have you looked at whether their container image allows installing arbitrary pip packages? That was another backdoor for us.
That `session.trust_env = False` bypass is exactly why we stopped relying on proxy envs for policy enforcement. It's a configuration, not a control.
The container image angle is critical. If the base image can `pip install`, your network layer is still secure, but you're introducing risk from untrusted code execution. We locked this by using a custom image built from a minimal base with a frozen `requirements.txt` and removing `pip` entirely. The build pipeline scans the spec for any new packages.
Your pre-execution scanner idea is the right layer. We run a simple AST parse on any new script before it hits the repo, looking for `import` statements and HTTP client instantiations. It flags anything not on our internal registry allowlist. It's a gate, not a guarantee, but it shifts the failure earlier.
Right-size or die
The frozen `requirements.txt` and removing `pip` is a smart move. It makes me wonder, have you benchmarked that approach against something like a managed SaaS tool for this, like Snyk or even GitHub's Dependabot? I've been looking at the runtime overhead.
The AST parser for pre-flight checks is clever. We tried something similar but found it missed dynamic imports like `__import__('requests')`. We ended up adding a linter rule to catch those patterns, but it's a constant cat-and-mouse game.
How do you handle scripts that pull in other scripts? Our scanner sometimes missed nested calls because it only looked at the entry point.
Benchmarking my way to better decisions
That "defined boundaries" line gets me every time. It's marketing fluff for "you build the sandbox, we sell the shovel." Your network layer plan is the only real control you've got.
I've seen the proxy environment variables get bypassed by a custom `requests.Session` or even a naive `urllib` call. The VPC lockdown will catch it, but you're relying on a network deny as your primary alert. That's reactive, not preventive.
The real question is why their agent runtime doesn't have a built-in, mandatory allow list for egress calls. It's 2025. Every other automation tool figured this out a decade ago.
You're stuck building the guardrails they omitted.
Trust but verify
Your layered plan is correct. The proxy config alone isn't enough, as others have noted.
Key caveat: the `http_proxy` environment variables are useless if the agent's script uses a library that ignores them. You need to enforce it at the network layer with that default-deny egress, making the VPC your single source of truth.
Also, log those proxy denies. If a script tries to call out, you need the alert to trigger an immediate block of that script version, not just a network error.
Five nines? Prove it.
That proxy config snippet is the exact trap I fell into. It only works if the script uses a library that reads those vars by default.
You mentioned the scanner looking for http:// strings. Add checks for any `requests.get/post` and `urllib.request.urlopen` calls. Our scanner also flags `socket.create_connection` on ports 80/443.
But yeah, network layer's the real backstop. Those script scans just give you a heads-up before the VPC block hits.
Optimize or die.
Absolutely, the script scan is just a tripwire. The VPC block is the only thing that actually stops the call. We treat any scanner hit as a deployment blocker, but it's amazing how many ways there are to hide a network call. We've even seen base64-encoded URLs fetched via `exec`.
You can't scan for everything, but you can make the network say no.
measure twice, ship once
Yeah, logging the proxy denies and linking them to the script version is smart. It turns a network event into an actionable audit trail.
But how do you actually link the two? If a call gets blocked at the VPC, you get an IP and timestamp. How do you trace that back to a specific agent or script version? Is that something your logging setup handles automatically?
Mandatory configuration files are a solid idea in theory, but I've found they rely on the agent runtime actually respecting them, which isn't always guaranteed. OPA is a robust policy engine, but the overhead of injecting it into every agent execution can become its own performance bottleneck.
Your rule about a pre-approved, version-controlled repo is the key control. We enforce this by integrating the scan into the PR process. Any script modification triggers a static analysis job that looks for network calls and undeclared dependencies. It's not perfect, but it shifts the failure left. The real trick is making the scan fast enough that developers don't work around it.