We trialed Fin for 3 months on a 10-agent support team. Ran it against our standard benchmark suite: 500 historical tickets, measured deflection rate, agent edit rate, and time-to-resolution.
Key results:
* **Deflection Rate:** Claimed 50%+, we saw 22.3% on first reply.
* **Agent Edit Required:** 68% of AI-generated replies needed significant editing before sending.
* **Complex Query Failures:** Broke on any ticket requiring database lookup or multi-step logic.
Here's our config and the main test loop:
```python
# Simplified test harness
def evaluate_fin_response(ticket):
ai_reply = fin.generate(ticket.body)
if validate_as_final(ai_reply):
return "deflected"
elif requires_minor_edit(ai_reply):
return "light_edit"
else:
return "heavy_edit"
# Results summary
benchmark_results = {
"total_tickets": 500,
"deflected": 112,
"light_edit": 48,
"heavy_edit": 340
}
```
At $39/agent/month, the ROI isn't there for us. It's fast for simple FAQ replies, but agents spend more time correcting than solving. For a team our size, you're paying ~$4.7k/year for a mostly unreliable first draft.
- bench_beast
Benchmarks don't lie.