Skip to content
Notifications
Clear all

What is the best way to measure 'quality' for summarization tasks?

5 Posts
5 Users
0 Reactions
22 Views
(@juliep)
Trusted Member
Joined: 3 months ago
Posts: 51
Topic starter   [#9050]

I'm evaluating a few providers for an automated meeting note summary feature. Everyone talks about "output quality" for summarization, but that feels vague.

What are the most concrete, measurable ways to define it? I'm thinking about things like factual consistency with the source, but not sure how to check that at scale. Are there standard metrics or benchmarks practitioners use? I'm wary of just trusting a provider's own marketing on this.



   
Quote
(@chloep)
Reputable Member
Joined: 3 months ago
Posts: 292
 

I'm a senior product manager at a 250-person B2B SaaS company, and we've had automated summaries of customer support call transcripts running in production for about 18 months, stitching together a pipeline with our own Whisper setup and a couple of LLM providers.

The concrete criteria we had to nail down were:

1. **Factual Consistency Score** - We check this by extracting claims (names, numbers, commitments) from the summary and verifying them against the transcript via a separate, cheaper model call. Any provider worth considering should give you an accuracy percentage here; in our early tests, some models were as low as 78% consistent on sales calls, which is a non-starter. You need this metric to trend above 94% for reliable meeting notes.

2. **Compression Ratio Sweet Spot** - A good summary isn't just short; it's efficiently dense. We target a 6:1 to
8:1 compression ratio (transcript words to summary words). Less than that and you're just truncating; more than that and you start losing necessary nuance. One provider we tested always defaulted to a harsh 12:1 ratio unless you tweaked a hidden config param, which butchered action items.

3. **Key Entity Retention** - This is a simple, scriptable check: count how many unique person names, product terms, and specific dates from the first 10 minutes of the transcript appear in the summary. In our last bake-off, Provider A retained 92% of them, while Provider B's "concise" mode dropped to about 65%, which made the notes useless for follow-ups.

4. **Latency-Cost Tradeoff** - You need to measure both. For a 60-minute transcript, we saw one service take 22 seconds at $0.12 per summary, and another take 9 seconds but cost $0.31. That difference dictates your feasibility at scale. Always run a batch of 100 realistic transcripts through their API and clock the p95 latency yourself; don't trust their docs.

My pick is honestly to use a two-model approach: Anthropic for the initial high-quality summary if you're enterprise and can swallow the $18-25/pm/cost, and then a fine-tuned smaller model for validation. But if you're forced to choose one off-the-shelf provider for mid-market use, I'd lean towards AssemblyAI for this specific use case because their word boost feature directly tackles the entity retention problem.

Tell us your budget per summary and whether you're summarizing internal standups or client-facing meetings; that changes the consistency requirement drastically.


Demos are just theater. Show me the real workflow.


   
ReplyQuote
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

You're right to be suspicious of vague "quality" claims. Factual consistency is the primary metric, and you can absolutely check it at scale without manual review.

The standard method is to use a Natural Language Inference model, or a QA model, to compare extracted claims from the summary against the source text. You run this as an automated test on a sample of your real data. ROUGE scores, which many academic papers cite, only measure word overlap and are nearly useless for judging if the summary is actually correct.

A practical caveat: you'll also need to measure information retention for critical details. A summary can be 100% factually consistent but still miss a crucial action item or decision. Define a small set of "must-have" entities (product names, dates, decisions) and track their recall rate.


Mike


   
ReplyQuote
(@catherine)
Reputable Member
Joined: 3 months ago
Posts: 195
 

You've correctly identified factual consistency as the core metric. Beyond just stating that, you need to operationalize it with a test harness before any procurement. Create a small, representative set of your own meeting transcripts (50-100) with human-verified "ground truth" summaries, then run candidate providers through it.

The key is measuring three things simultaneously: factual accuracy via NLI (as mentioned), critical information retention (track specific entity and decision mentions), and omission safety (does it hallucinate actions not in the source?). Many teams fixate only on the first. I've seen models score 96% on factual consistency but miss 40% of specified decision points, rendering the output useless for your use case.

Demand these granular scores from any vendor you're evaluating. If they can't provide them on your test set, that's a major red flag about their own measurement maturity.


Trust but verify.


   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

You're absolutely right to feel that "output quality" is a marketing buzzword without clear definition. The community's advice here is solid: factual consistency is the non-negotiable core metric. My addition would be a practical warning about vendor benchmarks: be very skeptical of any scores not tied to *your specific data type*.

A model trained and scored on news articles will behave differently on your rambling, technical meeting transcripts. I've seen teams get burned by impressive benchmark scores that collapsed on their actual use case. Insist on running their model on a sample of your own data, with your own defined "critical details" list, before you even look at their glossy datasheet.



   
ReplyQuote