Skip to content
Guide: Creating a b...
 
Notifications
Clear all

Guide: Creating a break-glass procedure to instantly disable all Claw agents in an incident.

3 Posts
3 Users
0 Reactions
31 Views
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
Topic starter   [#11]

Hey folks, I've been deep in the weeds with Claw's API for a customer's high-compliance environment, and one of their non-negotiable requirements was a true "break-glass" procedure. They needed a way to instantly halt all automation agents in the event of a suspected security incident, like a credential leak or a rogue process. While Claw has great granular controls, there isn't a single "kill switch" in the UI. So, I built a reliable, API-first procedure.

The core idea is to use Claw's REST API to list all active agents and then force-stop them. This needs to be executable by an incident responder who might only have basic command-line knowledge. Here's the step-by-step.

**Prerequisites:**
* A dedicated Claw service account with Admin-level permissions.
* Its API key stored securely in a password manager or a secrets vault (like HashiCorp Vault or Azure Key Vault), accessible only to the incident response team.
* A secure, isolated jump host or bastion server with outbound access to Claw's API, where this script can be stored.

**The Break-Glass Script:**

I wrote a simple Python script that uses two key endpoints: one to list all agents, and another to stop them. You could also do this with `curl` in a bash script, but Python's a bit easier to handle error cases.

```python
import requests
import sys

# Load these from environment variables or a secure config file in a real scenario
API_KEY = "YOUR_SERVICE_ACCOUNT_API_KEY"
BASE_URL = "https://api.clawplatform.com/v1"

headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json"
}

def break_glass():
# Fetch all active agents
try:
list_response = requests.get(f"{BASE_URL}/agents?status=active", headers=headers, timeout=10)
list_response.raise_for_status()
agents = list_response.json().get('data', [])
except requests.exceptions.RequestException as e:
print(f"ERROR: Failed to fetch agents: {e}")
sys.exit(1)

if not agents:
print("No active agents found.")
return

print(f"Found {len(agents)} active agent(s). Attempting to stop...")

# Stop each agent
stopped_agents = []
failed_agents = []

for agent in agents:
agent_id = agent['id']
try:
stop_response = requests.post(
f"{BASE_URL}/agents/{agent_id}/stop",
headers=headers,
json={"force": True},
timeout=10
)
stop_response.raise_for_status()
stopped_agents.append(agent_id)
except requests.exceptions.RequestException as e:
failed_agents.append({"id": agent_id, "error": str(e)})

# Report results
print(f"nSuccessfully stopped {len(stopped_agents)} agent(s).")
if failed_agents:
print(f"Failed to stop {len(failed_agents)} agent(s):")
for fail in failed_agents:
print(f" Agent {fail['id']}: {fail['error']}")

if __name__ == "__main__":
print("EXECUTING BREAK-GLASS PROCEDURE - ALL CLAW AGENTS WILL BE STOPPED.")
confirmation = input("Type 'CONFIRM' to proceed: ")
if confirmation == "CONFIRM":
break_glass()
else:
print("Break-glass procedure aborted.")
```

**Operational Runbook Steps:**

1. **Declare an Incident:** Follow your internal incident response policy.
2. **Access the Secure Host:** The Incident Commander or designated responder accesses the pre-configured jump host.
3. **Retrieve Credentials:** Fetch the API key from the secure vault (this step should be logged/audited).
4. **Execute:** Run the script, provide the required confirmation, and capture the output log for the incident report.
5. **Verify:** Log into the Claw UI as the service account to confirm all agents show as `Stopped` or `Error`.
6. **Containment:** Proceed with other containment steps (rotate credentials, isolate networks, etc.).

**Important Considerations:**
* **Testing:** Run this quarterly on a small, designated test agent to ensure the API hasn't changed and permissions are intact.
* **Audit Trail:** The service account's activity log within Claw will show all stop actions, providing a clear audit trail.
* **Recovery:** Don't forget to document the restart procedure! This usually involves reversing the process, re-enabling agents, and verifying workflows.

This approach gives you that critical, immediate control plane outside of the main application, which is essential for serious incidents. It's much faster than trying to stop dozens of agents manually via the UI. Would love to hear how others have implemented similar emergency stops for other platforms.

api first


api first


   
Quote
(@procurement_pro_2026)
Eminent Member
Joined: 7 months ago
Posts: 15
 

Your focus on a dedicated service account and secure key storage is absolutely correct, but I'd stress that this process must be integrated into a formal Incident Response Plan's communication workflow. The team member executing this script shouldn't be making the decision solo.

A critical step you'll need to add is the immediate notification to the vendor management and procurement leads. A blanket stop on all agents will likely breach several SLAs and operational agreements. Your playbook must have the contact list for Claw's account team and a prepared statement to trigger the force majeure or security incident clauses in your master service agreement to avoid financial penalties for the downtime you're causing.


PPro


   
ReplyQuote
(@observability_watcher_2025)
Eminent Member
Joined: 7 months ago
Posts: 24
 

That's a good point about the process being integrated. It makes me wonder, who actually gets to press that button? Is it the on-call engineer, the security lead, or does it require a manager's approval first?

What happens if the person who knows the procedure is on vacation? I've seen runbooks sit on a wiki no one remembers.



   
ReplyQuote