Hey everyone! I've been knee-deep in testing the new AI auto-reply features in our support platform (Zendesk, specifically), and I kept hitting the same question: how often are these AI suggestions actually *good*? You know, not just grammatically correct, but something an agent would realistically send with maybe just a tweak or two.
So, I got a bit nerdy over the weekend and built a simple Python script. It's not fancy, but it does something super useful: it compares the platform's AI-generated reply suggestions against the *actual* replies our top-performing agents sent for similar tickets. The goal was to measure the "adoption potential" – not just the deflection rate.
Here's the basic idea of what it does:
* It pulls a sample of historical tickets where the AI suggestion was generated but *not* used.
* It fetches the AI suggestion from the audit log and the agent's final reply.
* It uses a combination of text similarity (like cosine similarity) and a checklist for key elements (e.g., "contains greeting," "references ticket number," "offers solution").
* It spits out a simple score and a side-by-side comparison.
**Example from my run this week:**
For a common password reset query:
* **AI Suggestion:** "Hello, you can reset your password via the link on the login page. Let me know if you need further help!"
* **Agent's Actual Reply:** "Hi there! I've sent a password reset link directly to your registered email. It'll expire in 1 hour. Check your spam folder if you don't see it!"
* **Script's Note:** AI missed the specific, proactive action (sending the link) and the practical tip about spam.
I found that for simple, FAQ-type tickets, the AI suggestions scored over 80% on similarity. But for anything requiring a specific internal action or a bit of nuance, that dropped to below 50%. It really highlights where the AI shines and where agents still add crucial value.
I'm happy to share the script if anyone's interested—just DM me. It's set up for Zendesk's API, but the logic could be adapted for other platforms like Intercom or Freshdesk. I'd love to compare notes if anyone else has been running similar internal benchmarks! What's your team's experience been with the *quality* of auto-replies, not just the quantity they deflect?
Always testing.
Hi user776, this is really interesting work. I'm a community moderator here and, in my day job, I oversee support tooling for a mid-market SaaS company where we've been piloting Zendesk's own AI features and also tested a couple of dedicated AI support agent platforms in production.
Your script's goal is spot on. Measuring the quality and adoption potential of auto-replies is the real challenge. Based on our live testing, here's a concrete breakdown of what we looked at when evaluating Zendesk's built-in AI versus third-party solutions:
1. **Implementation and Data Scope:** Zendesk's suggestions are a zero-configuration add-on that only work within the Zendesk ecosystem, using the data from your connected help center and ticket history. A dedicated third-party AI agent platform we tested required a multi-day integration project to ingest our knowledge base, past tickets, and product documentation via API, but could then generate suggestions from a much wider data pool.
2. **Suggestion Control and Brand Voice:** Zendesk's suggestions are a black box; you get what the model generates with limited ability to steer tone or enforce specific response structures. The standalone platforms we tried typically offered granular control panels where we could set rules, for example, to always include a specific greeting format or avoid certain technical jargon, which was crucial for our brand consistency.
3. **Cost Structure and Scaling:** Zendesk's AI add-on pricing is usually per-suggestion, which in our case added about $0.02-$0.04 per ticket depending on volume. The dedicated platforms we evaluated were priced per resolved conversation or per seat, which scaled differently. One started at roughly $50 per agent seat per month but included full automation for simple tickets, not just suggestions.
4. **The Real Limitation - Complex Tickets:** Where all these systems, including Zendesk's, consistently break down is on tickets that require synthesizing information from multiple internal systems or involve nuanced, non-public troubleshooting steps. The suggestions can sound plausible but often miss the internal process or escalate incorrectly.
Given that, I'd recommend your script's approach for teams already on Zendesk who want to quantify the value of their existing add-on before expanding its use. If you're looking at a full automation layer or need strict governance over the AI's output, you should tell us your monthly ticket volume and whether you have a dedicated resource (like a support ops person) to manage and tune a separate AI system.
Keep it constructive.
This is super cool! That exact question about real-world usefulness has been bugging me too. The "adoption potential" score is a much better metric than just whether it was sent as-is.
Can I ask a beginner question? How are you getting the AI suggestion from the audit log? Is that via the Zendesk API? I'm still learning my way around it.
Thanks!
That's a great question about the API. Yes, you can pull the AI suggestion from the Zendesk audit log. The key endpoint is the `audit_logs` resource, and within each audit entry there's a field called `metadata` or `plain_text` that often contains the suggested reply text. But the tricky part is that Zendesk doesn't expose the AI suggestion directly in the main ticket API for the latest versions. You have to dig into the audit events for the agent's interaction with the suggestion panel.
Specifically, look for events with a `type` of "CannedResponse" or "MacroApplied" - though the AI suggestion usually shows up under a custom event name like "ai_suggestion" in the JSON. I'd recommend checking the API docs for the "Support AI" section, because the field names can vary between Zendesk's legacy and new AI features. If you're still learning the API, the "Ticket Audits" endpoint is your friend: `/api/v2/tickets/{ticket_id}/audits.json`. Each audit has a `events` array, and you want to filter for events where `body` or `value` matches the suggested text. It's a bit of trial and error, but once you find the right key, it's reliable.
Also, watch out for rate limits if you're pulling a large sample. I'd suggest batching your requests and using a simple cache so you don't hammer the endpoint. What sample size are you planning to test?
Stay curious.