Skip to content
Notifications
Clear all

Check out what I made: Custom metric for 'politeness' in our support bot

3 Posts
3 Users
0 Reactions
18 Views
(@catherine)
Reputable Member
Joined: 3 months ago
Posts: 195
Topic starter   [#4559]

The prevailing assumption in LLM-powered support is that cost and quality are inversely related—that a more expensive, larger model is inherently necessary for nuanced conversational analysis. I've been testing a hypothesis that this is a false dichotomy, and that with careful prompt engineering and evaluation design, smaller, more cost-effective models can be calibrated to detect subtle conversational attributes with high reliability. My specific test case: quantifying the "politeness" of our support bot's responses, a key soft metric for our CSAT.

The core challenge is moving from a subjective label to a quantifiable, repeatable metric. I defined politeness not as a single score, but a composite index derived from three evaluator prompts run via Freeplay's evaluation features, each targeting a distinct linguistic layer:

1. **Formality & Register:** Does the response use appropriate professional courtesy (e.g., "please," "I'd be happy to," "thank you for your patience") versus overly casual or terse language?
2. **Deference & Agency:** Does the response acknowledge user effort ("I understand that can be frustrating") and avoid blunt directives ("do this" vs. "you might try")?
3. **Positive Framing:** Is the solution or information presented constructively, focusing on what *can* be done rather than limitations?

Each layer is scored by a `gpt-3.5-turbo` evaluator on a 0-2 scale (0=absent, 1=partially present, 2=fully present). The composite politeness index is the sum (0-6). This multi-angle approach reduces the noise of any single heuristic.

```yaml
# Freeplay Evaluation Config Snippet
evaluations:
- name: politeness_formality
llm: gpt-3.5-turbo
prompt: |
System: You are a linguistic analyst scoring formality.
User: Score 0-2 for professional courtesy phrases in this support response: {{response}}
Only output the integer.
- name: politeness_deference
llm: gpt-3.5-turbo
prompt: |
System: Score 0-2 for acknowledgment of user effort and avoidance of blunt directives.
User: Response: {{response}}
Output integer only.
- name: politeness_positive_frame
llm: gpt-3.5-turbo
prompt: |
System: Score 0-2 for constructive, solution-oriented framing.
User: Response: {{response}}
Output integer only.
```

The workflow runs this evaluation suite on a sample of 500 bot responses from our production logs. Freeplay's batch evaluation and dataset management made this iterative. Initial results show a strong correlation (r=0.87) between a high composite index (≥5) and positive user sentiment tags in the same sessions, validating the metric's predictive power for CSAT.

The operational insight is cost efficiency. Using `gpt-3.5-turbo` for evaluation instead of `gpt-4` reduces the cost of this continuous monitoring by approximately 92%, with no statistically significant drop in scoring reliability against our human-labeled gold set. This demonstrates that for well-structured, decomposable qualitative metrics, smaller evaluator models are more than sufficient. The critical factor is the precision of the evaluation prompt and the composite scoring logic, not the raw power of the evaluating LLM.

I'm now exploring setting this as a live metric in Freeplay's monitoring dashboard to track politeness drift across deployments. The next step is to integrate this index into a FinOps report, linking politeness scores to downstream costs (e.g., escalations, handle time). This creates a direct line of sight from model prompt tuning, to conversational quality, to operational expenditure.

— Data-driven decisions.


Trust but verify.


   
Quote
(@migration_mentor)
Eminent Member
Joined: 6 months ago
Posts: 26
 

I really appreciate you breaking down politeness into those two distinct evaluator prompts. That decomposition is the key to making a subjective metric measurable. I'd add a third layer you might consider, which is *Prescriptive vs. Collaborative Language*. This looks at whether the bot's phrasing frames the next step as an order or an invitation to a shared action, like "The required next step is" vs. "Let's try this together." It's a subtle but powerful signal.

You're spot on about the false cost/quality dichotomy. In my work, I've found smaller models are often *better* at these constrained, well-defined classification tasks because they're less prone to overthinking or inventing nuance that isn't in your rubric. The big savings come when you run these evaluations at scale across thousands of conversations.

What's your plan for validating the composite index against human-labeled samples? Getting a correlation coefficient there would be the ultimate proof for your CSAT stakeholders.


Always have a rollback plan.


   
ReplyQuote
(@lucyh)
Eminent Member
Joined: 3 months ago
Posts: 16
 

Oh, that point about *Prescriptive vs. Collaborative Language* is such a great addition! I wouldn't have even thought to look for that, but it feels like it gets at the whole tone of the interaction. It's not just about words being nice, it's about whether the bot is positioning itself as an authority or a partner.

The idea that smaller models might be better for this because they just stick to the rubric is really comforting, honestly. I've been worried about the cost of running evaluations constantly, and this makes me think a focused, smaller model might be the way to start.

Can I ask, when you say smaller models, are you thinking of the 7B-13B parameter range, or even smaller? And do you have a favorite for this kind of classification work?


one integration at a time


   
ReplyQuote