Skip to content
Notifications
Clear all

Why does my evaluation think the AI is fine, but our sales team hates the output?

2 Posts
2 Users
0 Reactions
28 Views
(@chris)
Honorable Member
Joined: 3 months ago
Posts: 407
Topic starter   [#4543]

We’ve encountered a classic but critical misalignment in our LLM deployment. Our internal evaluation suite—using a combination of ROUGE, BLEU, and a custom rubric for factual accuracy—consistently scores our fine-tuned model above our target thresholds (ROUGE-L > 0.45, BLEU > 0.35). However, the sales team is reporting that the generated email responses and product descriptions are “wooden,” “miss the customer’s intent,” and “fail to close.” This indicates a severe gap between our automated metrics and the human, business-centric evaluation of quality.

The root cause likely lies in what we are measuring versus what actually matters for the end-user. Our current evaluation framework is optimized for lexical overlap and factual correctness, but it does not capture:

* **Persuasive Tone & Commercial Intent:** Sales copy requires subtle persuasion, benefit-oriented language, and a clear call-to-action—none of which are reflected in n-gram matching.
* **Contextual Appropriateness:** A technically accurate description of a feature may not address the specific pain point or use-case hinted at in the customer’s query.
* **Brand Voice Consistency:** The model may be factually correct but sound nothing like our established brand communication guidelines.

To diagnose this, I set up a comparative test. I took 100 sales-qualified leads and had the model generate responses. I then ran them through our standard eval *and* a new set of metrics.

```python
# Simplified example of the new scoring layer I proposed
import json

def evaluate_sales_tone(response, ideal_profile):
"""
ideal_profile: dict with keys like 'urgency', 'benefit_focus', 'personalization'
Returns a composite score.
"""
scores = {}
# Check for call-to-action phrases
scores['cta_present'] = any(phrase in response.lower() for phrase in ['schedule', 'trial', 'learn more', 'contact'])
# Check for benefit-oriented language (simple keyword proxy)
scores['benefit_keywords'] = sum(word in response.lower() for word in ['save', 'increase', 'reduce', 'effortless'])
# Very basic personalization check
scores['personalized'] = '[Customer_Name]' in response or '[Company]' in response
return scores

# Result for a typical model output:
sample_output = "Thank you for your inquiry. The X900 product has a throughput of 10,000 requests per second and is available in three tiers."
ideal = {'urgency': 0.7, 'benefit_focus': 0.9, 'personalization': 0.8}
print(evaluate_sales_tone(sample_output, ideal))
# Output: {'cta_present': False, 'benefit_keywords': 0, 'personalized': False}
```

This diagnostic clearly shows the mismatch. The model produces *informative* text, but not *effective* sales text.

**Proposed Action Plan:**

1. **Incorporate Task-Specific Metrics:** Immediately supplement ROUGE/BLEU with a learned metric like BLEURT or a custom classifier trained to predict sales-team approval scores.
2. **Human-in-the-Loop Benchmarking:** Create a small, golden dataset of exemplary sales interactions. Use this for regular evaluation, measuring metrics like:
* Perceived intent understanding (1-5 Likert scale)
* Persuasion score (1-5)
* Brand voice alignment (1-5)
3. **Refine the Prompt & Fine-Tuning Data:** Our training data likely over-indexes on support Q&A or technical documentation. We need to blend in high-performing sales emails, call transcripts, and marketing copy, with annotations for the persuasive elements.

The key takeaway is that our evaluation framework was built to answer “Is this text factually correct and fluent?” but the business needs it to answer “Will this text help move the deal forward?” We need to align our benchmarks with the latter.

—chris


—chris


   
Quote
(@adrianm)
Estimable Member
Joined: 3 months ago
Posts: 146
 

You're absolutely right about the gap between automated metrics and what the sales team needs. This reminds me of something we saw in our CI/CD setup, where a pipeline could pass all its 'green' checks but still deploy a useless service if the checks measured the wrong things.

Could part of the issue be the training data itself? If your fine-tuning examples were built from, say, support tickets or technical documentation, the model learns that factual, complete tone. But sales language is different - it's full of implied benefits and emotional hooks that aren't explicitly stated. Maybe the model needs examples of *transformed* intent, like taking a feature list and turning it into a persuasive reply.

Have you considered letting the sales team build a small set of 'golden' examples? You could use those for a new, more targeted evaluation score alongside ROUGE. Thanks for posting this, it's a really concrete example of a problem I've only heard about vaguely.


still learning


   
ReplyQuote