Skip to content
Showcase: My simple...
 
Notifications
Clear all

Showcase: My simple scorecard for weekly agent performance review with the team.

1 Posts
1 Users
0 Reactions
6 Views
(@crm_hopper_2025)
Estimable Member
Joined: 2 months ago
Posts: 113
Topic starter   [#3977]

Alright, team. I know this is a bit outside my usual CRM migration war stories, but hear me out. I’ve been living in this weird space between RevOps and AI SOC for the last six months, and let me tell you, managing a team of AI agents feels eerily similar to onboarding a new sales team on a fresh CRM instance. You have high hopes, you feed it data, and then… you need to see if it’s actually doing its job or just creating elegant, hallucinated ticket art.

We’ve got a handful of specialized agents now—one for initial triage, one for deeper investigation, one for enrichment—and our weekly sync was turning into a "vibe check." Super subjective. "Agent B felt slow this week." Not helpful. So, I built a dead-simple scorecard we review every Monday. It’s not fancy, but it’s forced us to look at concrete outcomes, not just feelings. It’s basically the equivalent of checking your data mapping logs after a big migration—painful but necessary.

Here’s what’s on the card for each agent. We pull this from our SIEM, ticketing system, and the agents' own audit logs:

* **Volume Handled:** Raw number of alerts/events ingested. (The "how busy were you?" metric).
* **Deflection Rate:** Percentage of items resolved without human escalation. This is our big efficiency win.
* **Mean Time to Triage (MTT):** For the triage agent, how fast it sorts an alert into "ignore," "investigate," or "escalate."
* **Investigation Quality Score:** This one’s manual. We randomly sample 25 of its investigation summaries and grade them (1-5) on accuracy, relevance, and actionability. It’s time-consuming but stops complacency.
* **False Positive Creation:** How many new tickets did this agent open that we, as humans, closed within 5 minutes as nonsense? Keeps the "enthusiasm" in check.
* **Tool Call Success Rate:** Percentage of times its API calls to enrichment sources (VT, whois, etc.) succeed vs. error out. Integration health 101, people!

The magic isn't in any single number. It's in the trends and the conversations. Last week, Triage Agent's deflection rate dipped 5%. Why? We dug in and saw a new alert source was formatted weirdly, confusing its little logic. Without the scorecard, we might have just blamed "the AI" and moved on.

It’s a living doc. We already added a "Cost Per Resolution" column after realizing one agent was making crazy expensive API calls for minor alerts. Think of it like monitoring your Salesforce API limits or HubSpot workflow costs.

Does anyone else do something similar? I’m sure there are fancier LLM evaluation frameworks out there, but this operational, team-focused view has been a game-changer for us. Would love to hear what metrics you’re tracking, or what I’m painfully missing.

Hopefully last migration,



   
Quote