Skip to content
Notifications
Clear all

Guide: building a simple feedback loop to improve AI suggestions over time

3 Posts
3 Users
0 Reactions
43 Views
(@data_diver_42)
Honorable Member
Joined: 7 months ago
Posts: 400
Topic starter   [#3892]

Hey folks, been playing with the new AI-suggested replies in our support platform. The initial suggestions are... okay, but they often miss the mark on our specific product lingo or common edge cases. The real magic happens when you can actually *improve* them based on what your agents accept or edit.

I built a simple feedback loop to capture that signal and retrain the suggestion model (or at least give actionable data to the vendor). It's basically a three-step process:

1. **Log the interaction:** For every AI suggestion shown, log the suggestion ID, the ticket context (category, keywords), and the agent's final action.
2. **Score the outcome:** Did the agent use it as-is, edit it, or ignore it? We score a "use as-is" as a strong positive signal.
3. **Aggregate & analyze:** Weekly roll-up to see which suggestion patterns are working and which aren't.

Here's the core of the logging table schema I used:

```sql
CREATE TABLE ai_suggestion_feedback (
suggestion_id UUID PRIMARY KEY,
ticket_id INT,
suggested_text TEXT,
used_as_is BOOLEAN,
edited_text TEXT NULL,
final_text TEXT,
category VARCHAR(50),
feedback_ts TIMESTAMP
);
```

Then, a simple weekly query to find patterns in failing suggestions (e.g., by ticket category or keywords):

```sql
SELECT
category,
COUNT(*) as total_suggestions,
SUM(CASE WHEN used_as_is THEN 1 ELSE 0 END) as accepted_count,
AVG(CASE WHEN used_as_is THEN 1.0 ELSE 0.0 END) as acceptance_rate
FROM ai_suggestion_feedback
WHERE feedback_ts >= DATEADD(week, -1, GETDATE())
GROUP BY category
HAVING COUNT(*) > 10
ORDER BY acceptance_rate ASC;
```

This surfaces which categories have low acceptance, so you can either provide more tailored examples to your vendor or build a custom prompt for those cases. We saw a 15% bump in "as-is" usage for billing inquiries after feeding some examples from our best agents.

Has anyone else set up something similar? Curious if you're measuring the *edit distance* between the suggestion and the final reply as a more nuanced metric. Also, what's your storage look like for this kind of interaction data – are you piping it back to your data warehouse?

--diver


Data is the new oil - but it's usually crude.


   
Quote
(@martech_hopper_22)
Trusted Member
Joined: 6 months ago
Posts: 48
 

Nice, logging the data is the right start. But the scoring part is where most teams trip up.

"Use as-is" isn't always a perfect positive signal. Sometimes agents just hit send to close a ticket fast, even if the suggestion is mediocre. We also track edit *distance* - if they change two words vs. rewrite the whole thing, that's a huge difference in feedback quality.

Also, weekly analysis might be too slow if you're getting a lot of volume. We built a real-time dashboard flagging categories with a <20% "as-is" rate, so we can intervene and add specific training phrases.


Trial number 47 this year.


   
ReplyQuote
(@emilyk)
Reputable Member
Joined: 3 months ago
Posts: 286
 

The edit distance metric is crucial. We found that tracking character-level Levenshtein distance alone created false positives for simple punctuation changes, while ignoring semantic degradation.

We built a hybrid scoring system that weights:
- Edit distance as a percentage of original text length
- Presence of domain-specific term replacement
- Session timing data (suggestion-to-send latency under 2 seconds gets flagged for potential "fast-click" behavior)

This revealed that 40% of our "use as-is" entries were actually low-quality acceptances when cross-referenced with subsequent ticket reopen rates. The real-time dashboard approach you mention is sound, but you need to account for that noise floor before triggering retraining cycles.


Show me the numbers, not the roadmap.


   
ReplyQuote