Skip to content
Notifications
Clear all

How do I prevent Claw agents from calling external APIs we haven't vetted?

22 Posts
22 Users
0 Reactions
75 Views
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

Exactly. The default-deny egress is the only real enforcement.

But you've got to watch those flow logs. We had the rule, but the alerts weren't tuned. A burst of denies from a runaway script got lost in the noise until we saw the bill for the proxy load balancer.

Link the DENY events to the specific compute instance ID in your alert. Otherwise you're just watching a blinking light.


SQL is enough


   
ReplyQuote
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
 

Exactly. Flow logs are just raw data - the real trick is streaming them into something that can correlate denies with instance metadata in real time. We pipe VPC flow logs into a Kafka topic, then use a simple stream processor to join them with our instance lifecycle events.

If you're not aggregating those denies per script execution window, you'll miss the slow drip. A script making one call every 10 minutes might not trigger a threshold alert, but it's still doing something it shouldn't.

What's your aggregation window look like? We found 5 minutes was too long to catch fast bursts, but 30 seconds generated too much noise.



   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

You mention a NAT Gateway with strict rules. Wouldn't the NAT's egress IP itself show up in that referrer header you're worried about? Isn't that still an information leak, even if the call gets blocked later?



   
ReplyQuote
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
 

Your "defined boundaries" line from the vendor sums it up perfectly. They've outsourced the hard part of their security model to you.

The proxy environment variables are theater. As others have said, they only work if the underlying libraries respect them, which is a huge assumption. Your network layer plan is the only real control.

But I'll push back on one point: you're focusing on stopping the call, which is correct, but you mentioned the referrer header leak. That's the bigger immediate risk. Even with a NAT gateway, if a script makes a call before your network policy drops the packet, that header's already gone. You need to treat the agent's outbound interface itself as hostile. We run ours with a kernel module that strips the referrer header at the socket level, because you can't trust the application layer to behave.

The vendor telling you to "curate scripts" is a red flag. It means they know their sandbox is porous.



   
ReplyQuote
(@briana)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Oh, you're so right about checking for `requests.get` and `urllib.request.urlopen`. That's the first layer of our scanner, but you've reminded me of another sneaky one. We also scan for `subprocess.call` or `os.system` patterns that might invoke `curl` or `wget` from the shell. A dev once wrote a "helper" that did just that to bypass our Python lib checks.

And you're dead on - the network block is the backstop. Our scanner just gives us a chance to fail the build and ask "hey, what's this for?" before it even tries to hit the wire.


Backup first.


   
ReplyQuote
(@briang)
Estimable Member
Joined: 3 months ago
Posts: 119
 

>because you can't trust the application layer to behave.

That's a really good point about the header leak happening before the drop. I hadn't thought about the timing. Stripping at the socket level is clever, but that sounds like a deeper system-level change than I'm used to.

If we're stuck with the application layer, is there any effective way to control those headers at runtime? Maybe something that intercepts libc calls? It still feels like you're just hoping the agent doesn't find another way out.



   
ReplyQuote
(@dragonrider)
Honorable Member
Joined: 3 months ago
Posts: 367
 

Ugh, "curate our tooling scripts more carefully" is such a classic vendor cop-out. It translates to "you figure out the security model for our tool." Your layered approach is the only sane path forward.

I think your network layer plan is solid, but I'd push you to consider the agent container itself as the first boundary, before the VPC. We run ours with a read-only root filesystem and drop all capabilities except the bare minimum. It won't stop a determined script from trying a call, but it eliminates a whole class of "oops I installed curl" accidents. Pair that with your proxy env vars (even if they're just theater) and you've got a decent first container ring.

That referrer header leak you mentioned is the real nightmare fuel, though. Even if the call gets blocked at the NAT, the packet might already be out the interface with your internal IP in it. Have you looked at seccomp profiles or a sidecar proxy that mutates outbound headers? It's a pain, but it's the only way to be sure the header never gets set.


Try everything, keep what works.


   
ReplyQuote
Page 2 / 2