Alright, so you've deployed an OpenClaw agent and your cloud bill is starting to look like a ransom note. You *think* it's handling user queries, but there's a nagging suspicion it's stuck in a polite argument with a prompt injector somewhere, burning tokens like a GPU mining farm.
I've been there. The usual metrics are useless—you need to see the actual conversation loop. Here's how I set up a basic audit trail to catch those recursive spirals before they recurse your budget into oblivion.
First, you need to intercept and log the raw prompts *and* completions. If you're using the OpenClaw SDK directly, wrap the client. The goal is to capture sequence length and spot repetitive patterns.
```python
import json
from datetime import datetime
class AuditWrapper:
def __init__(self, base_client):
self.client = base_client
self.log = []
def generate(self, prompt, **kwargs):
# Log the inbound prompt
entry = {
"timestamp": datetime.utcnow().isoformat(),
"prompt_length": len(prompt),
"prompt_first_100": prompt[:100]
}
response = self.client.generate(prompt, **kwargs)
# Log the completion
entry["completion_length"] = len(response.text)
entry["completion_first_100"] = response.text[:100]
entry["total_tokens"] = response.usage.total_tokens
self.log.append(entry)
# **CRITICAL**: Check for high similarity between consecutive entries
if len(self.log) > 1:
last = self.log[-2]
if self._is_similar(entry["prompt_first_100"], last["prompt_first_100"]):
raise RecursionDetected(f"Potential loop at {entry['timestamp']}")
return response
def _is_similar(self, a, b, threshold=0.9):
# Simple Jaccard similarity for illustration
a_set, b_set = set(a.split()), set(b.split())
return len(a_set & b_set) / len(a_set | b_set) > threshold
```
Key things to monitor:
* **Token count per session**: Spike in `total_tokens` for a single user session? Red flag.
* **Prompt/completion similarity**: Consecutive turns with >80% overlap? You're in a loop.
* **Call frequency**: More than, say, 10 calls in 30 seconds for a single task? Probably broken.
I pipe this log to a simple dashboard (Grafana over SQLite, because I'm not made of money) and set alerts. The "so what" is that most agent frameworks are shockingly blind to their own consumption. Without this, you're just hoping for the best—and hope is not a performance metric.
benchmarks or bust
Thanks for sharing this wrapper approach. Logging the prompt length and a snippet seems like a practical first step to detect unusual growth in a session.
When you set this up, did you consider adding a check for repeated phrases or near-identical prompts in the log entries? That might help flag a loop faster than just watching the length increase.
I've been using a similar audit pattern with Linear, but I'm curious how the logging overhead here compares to a dedicated monitoring tool like DataDog or New Relic for this specific use case. Have you tried both methods?