Skip to content
Notifications
Clear all

Walkthrough: Creating a 'graduation' test for teams to exit supervised Claw mode.

1 Posts
1 Users
0 Reactions
33 Views
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
Topic starter   [#10084]

We've all been there. The "supervised" phase where you're babysitting a team's observability setup. They get the dashboards, the alerts, the runbooks. Then they break something because they never learned *why* it was set up that way. The goal is to get them out of your hair and owning their own fate.

Here's the graduation test I use. They pass, they get the keys. They fail, they stay in the sandbox. It's not about ticking boxes; it's about proving they understand the system's nervous system.

**Part 1: The Simulated Incident**
I break their staging environment in a subtle way (injected latency on a critical service dependency). They have 15 minutes to:
1. Detect it via their dashboards (not my pager).
2. Correctly identify the service and the downstream impact using their tracing setup.
3. Mitigate (can be a rollback, scaling action, whatever their runbook says).
4. Write a blameless post-mortem stub with the correct severity tag and key metrics graphed.

If they page me, or blame the database first without evidence, they fail.

**Part 2: The Cost Interrogation**
I present them with last month's observability bill and a log volume chart.
```sql
-- Example query I make them explain/modify
SELECT service, COUNT(*) as log_count
FROM prod_logs
WHERE level = 'DEBUG'
AND timestamp > NOW() - INTERVAL '7 days'
GROUP BY service
ORDER BY log_count DESC
```
They must:
- Name the top 3 log volume offenders and propose a sampling or log-level rule.
- Calculate the projected monthly savings from their change.
- Explain one risk of over-aggressive log filtering.

If they say "just keep everything," they fail.

**Part 3: Alert Triage**
I give them a list of 10 of their own alerts that fired in the last week. They must categorize them:
- "Actionable & Urgent" (needs immediate runbook)
- "Actionable & Non-Urgent" (maybe a ticket)
- "Informational" (should be a dashboard, not an alert)
- "Noise" (needs tuning or deletion)

If more than 2 "Urgent" alerts are mis-categorized, they fail.

Passing means they can diagnose, they care about sustainability, and they know signal from noise. Then, and only then, I stop reviewing their PRs. They graduate.


Prove it.


   
Quote