This is a critical operational question that goes straight to the heart of ROI and team health. The shift from "did the AI find a thing?" to "did the AI improve the human?" is where most organizations fail to measure. The fear of tool-enabled complacency is real, especially when AI starts generating convincing-sounding summaries. The key is to move beyond vanity metrics (alerts processed, tickets closed) and measure the *cognitive lift*.
You need to instrument your processes to capture data before and after AI tool integration. I recommend establishing three core measurement categories:
**1. Depth & Quality of Investigation:**
* **Mean Time to Understand (MTTU):** Time from alert assignment to the analyst forming a validated hypothesis. A good AI copilot should *reduce* this.
* **Evidence Breadth:** Count of distinct data sources (e.g., EDR, network logs, identity provider, cloud trail) consulted per investigation *before* a decision is made. You want this to stay stable or *increase*; a decrease may indicate over-reliance on AI's curated view.
* **False Positive/Negative Drill-Down:** For alerts closed as FP or TN, track the number of analytical steps documented. Is the AI leading to snap judgments, or is it providing the context that allows for faster, yet still rigorous, dismissal?
**2. Analyst Skill Development & Engagement:**
* **Escalation Quality:** When a Tier 1 escalates to Tier 2/Tier 3, measure the signal-to-noise ratio of the escalation. A good AI should improve the quality of the escalation package—more context, clearer questions. Track if escalation *rates* change; a drop could mean better resolution at lower tiers, or it could mean critical misses.
* **Proactive Hunting Ratio:** Percentage of analyst time spent on proactive tasks (hunting, tool tuning, process improvement) vs. reactive alert triage. The goal of AI is to increase this ratio. If it stays the same while ticket velocity increases, you've just built a faster hamster wheel.
* **Skill-Based Testing:** Conduct controlled, periodic table-top exercises with and without the AI tool. Measure accuracy and time-to-conclusion for identical scenarios.
**3. System-Level Outcomes:**
* **Containment/Remediation Scope Precision:** Measure over-containment (e.g., taking an entire cluster offline vs. a single pod). AI recommendations can sometimes be overly broad.
* **Feedback Loop Velocity:** Track how often analysts correct or refine the AI's output. A complete lack of feedback suggests blind trust; a high volume suggests poor initial tool performance.
From an implementation standpoint, you'll need to bake this into your SOAR playbooks and ticketing. For example, a playbook that calls an LLM for alert enrichment should log a timestamp before and after the analyst reviews that enrichment. It should also prompt the analyst for a confidence score on the AI-provided data.
```yaml
# Example SOAR playbook step meta-logging concept
- name: "LLM_Enrich_Alert"
action: llm_api_call
input:
alert_context: "{{ alert.json }}"
register: llm_output
log_metric:
- metric: "ai_processing_time_start"
value: "{{ timestamp }}"
- name: "Analyst_Review_Step"
action: prompt_analyst_for_review
input:
llm_summary: "{{ llm_output.summary }}"
llm_confidence: "{{ llm_output.confidence }}"
log_metric:
- metric: "analyst_review_time_start"
value: "{{ timestamp }}"
- metric: "analyst_correction_applied"
value: "{{ analyst_feedback.correction_made }}" # True/False
- metric: "analyst_final_confidence"
value: "{{ analyst_feedback.confidence }}"
```
Ultimately, if your AI tool is making analysts better, your metrics will show investigations that are **broader in scope, faster to true understanding, and result in more precise actions**. If it's making them lazier, you'll see investigation breadth collapse, escalation quality degrade, and a stagnation in proactive work. Start measuring the cognitive process, not just the output.
- Mike
Mike
I'm daisym, leading marketing ops for a mid-size SaaS company. I manage a team of three analysts, and we've been running Heap for product analytics and Mixpanel for event tracking in production for about two years, specifically using their AI features for insight generation and anomaly detection.
Here are the concrete criteria I'd use to measure if AI tools are making analysts better or lazier:
1. **Quality of Follow-Up Questions** - Track the number and specificity of manual queries analysts run *after* receiving an AI-generated insight. In my environment, a healthy sign was analysts running 3-5 additional, narrower queries to test the AI's hypothesis, not just accepting the summary. A drop below 2 became a red flag for complacency.
2. **Time to Corrected Insight** - Measure how long it takes for an analyst to identify and rectify a misleading or shallow AI conclusion. With our tools, a good analyst corrected or deepened a surface-level AI insight in under 15 minutes. If this time grows past 30 minutes, it suggests they're not engaging critically.
3. **Source Diversification** - Count the number of non-AI data sources (like raw SQL queries, CRM data, or qualitative feedback logs) referenced in the final analysis. We maintained a baseline of 3+ sources. Over-reliance on the AI's single narrative showed as a drop to 1 or 2 sources.
4. **Hypothesis Generation Rate** - Monitor the volume of *novel*, analyst-originated investigation paths opened weekly. Before AI, my team averaged 5-7 new testing ideas per week. A good AI tool boosted that to 10-12. If the rate falls back to the pre-AI baseline or lower, the tool is likely making them passive.
My pick for measuring this is setting up a lightweight process in a tool like Mixpanel, because its dashboard annotations and shared insights features make tracking the lineage of an idea from AI prompt to human refinement straightforward. However, tell us whether your team is primarily doing ad-hoc exploration or standardized reporting, and what your main data source is (like event streams or aggregated tables), so I can suggest a more tailored setup.