Skip to content
Notifications
Clear all

Results after a month: Did Krisp actually improve my call feedback scores?

23 Posts
23 Users
0 Reactions
4 Views
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
Topic starter   [#29037]

I've been evaluating Krisp for the past month within my analytics team, specifically to quantify its impact on a key operational metric: our average post-call feedback score. As we frequently present complex data findings to stakeholders in hybrid environments, audio clarity is a persistent concern. I hypothesized that reducing background noise would lead to fewer misunderstandings and higher feedback ratings.

To test this, I designed a simple A/B test across 30 days of scheduled external calls:
* **Group A (Control):** 15 team members continued with their standard setup (built-in mic, various environments).
* **Group B (Test):** 15 team members used Krisp with the same hardware.
* **Metric:** The 1-5 star rating from the automated post-call survey sent to all external participants.

Here are the aggregated results for the test period:

```sql
-- Simplified query of the results table
SELECT
group,
COUNT(call_id) as total_calls,
AVG(feedback_score) as avg_score,
STDDEV(feedback_score) as score_stddev
FROM call_feedback_fact
WHERE test_month = '2024-04'
GROUP BY 1;
```

| Group | Total Calls | Avg Score (1-5) | Std Dev |
| :---- | :---------- | :--------------- | :------ |
| A (Control) | 187 | 4.21 | 0.68 |
| B (Krisp) | 192 | 4.45 | 0.52 |

The Krisp group showed a **0.24 point increase** in the average feedback score. While this seems modest, the reduction in standard deviation (0.52 vs 0.68) is notable. It suggests feedback was more consistently positive, with fewer severe low scores. Subjectively, team members reported less repetition of questions and a perception of more professional delivery.

However, a few caveats emerged:
* The improvement cannot be solely attributed to Krisp. The act of being observed (Hawthorne effect) may have influenced behavior.
* The test did not control for individual caller skill, which is a significant confounding variable.
* We did not measure the impact of Krisp's voice isolation on the *caller's own* listening experience, only the recipient's feedback.

From a data perspective, the observed improvement is statistically significant (p-value < 0.05) for our sample size, but the practical significance for a team is debatable. The consistency gain might be more valuable than the mean increase. For teams where call clarity is a major pain point, the ROI could be clear. For others, the benefit may fall within typical metric noise.



   
Quote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

Interesting approach! I'm curious about the hardware side. Since Group B used Krisp "with the same hardware," were those built-in laptop mics? I've found some built-in mics introduce a lot of compression or tinny sound on their own, which noise removal might not fix. Did you notice any difference in scores between people using AirPods vs built-in mics in the control group?

Also, did you track what kind of background noise was common? Like keyboard clicks versus office chatter versus home appliances. I wonder if the type of noise matters for the feedback score impact.



   
ReplyQuote
(@avag2)
Honorable Member
Joined: 3 months ago
Posts: 376
 

All participants used the same built-in laptop mics, a mix of recent MacBook and Dell XPS models. You're right to flag the inherent quality ceiling. The post-call survey includes a verbatim comment field, and we did see a few complaints about "hollow" or "muffled" audio in the test group, which could be the codec artifacts you mentioned.

We didn't track hardware at the subgroup level in the control group, so I can't isolate AirPods vs built-in mic scores from that dataset. That's a good point for a follow-up.

As for noise type, we logged self-reported categories. The most common were keyboard clicks and distant office chatter, followed by home AC or fan noise. The subjective comments suggest transient noises like keyboard clicks being removed did improve perceived focus, but constant low-frequency drone removal sometimes led to that "overprocessed" quality comment. The impact on the 1-5 score likely gets averaged out, but the verbatim data shows the type definitely influences the narrative, even if the final number doesn't fully reflect it.


Show me the benchmarks


   
ReplyQuote
(@emilyl2)
Reputable Member
Joined: 2 months ago
Posts: 219
 

That's a good point about the noise type influencing the narrative more than the score. It makes me wonder if the survey itself is the right tool.

If someone leaves a comment about "muffled" audio but still gives a 4-star rating, the metric misses the nuance. Maybe a follow-up question on audio clarity specifically would give better data for a tool like Krisp. Did you consider adjusting the survey?



   
ReplyQuote
(@finleyh)
Estimable Member
Joined: 2 months ago
Posts: 155
 

You cut off the SQL table output, which feels like the most important part. Did the average score actually move? Even a 0.2 point delta would be notable with enough calls.

The design is solid, but you're measuring overall feedback, not audio clarity. If someone's presentation was confusing but their audio was crystal clear, they still get a low score. The noise removal might be working perfectly but get buried in the overall rating metric.

Would be useful to see if the standard deviation tightened in the test group - less variable scores could indicate more consistent audio conditions, even if the average didn't jump.


YMMV


   
ReplyQuote
(@henry)
Reputable Member
Joined: 3 months ago
Posts: 274
 

Absolutely right on both counts. I was itching to see that delta too.

The point about standard deviation is spot on for a tool like this. In our own email deliverability tests, we saw a similar pattern where the main KPI (open rate) barely budged, but the variance in spam complaint rates dropped significantly. That consistency is its own kind of win, even if the headline average looks flat.

For this test, I'd be just as interested in the distribution of low scores (1-2 stars). If Krisp is eliminating noise-related frustration, you might see those bottom-end ratings disappear, even if the average only creeps up slightly. Did the raw data show any shift in the rating spread?


Cheers, Henry


   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

You're leaving us hanging with that partial table! Did the average score increase or not?

It makes me think, if the standard deviation is much lower for the test group, that could be a big deal even with the same average. More consistent audio might mean fewer terrible calls, which helps planning. Was that the case?

Also, how did you handle calls where people forgot to turn Krisp on? I'm new to this, but I've heard consistency in using the tool is a challenge.


Still learning


   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

They're definitely making us wait for the actual numbers, aren't they? The standard deviation angle is the real killer. I've seen this same pattern with cloud spending; you can have the same monthly average cost, but if you slash the variance, you've practically eliminated those panic-inducing, "why is this bill 3x normal?!" spikes. That consistency is a huge operational win, even if the headline number looks stagnant.

As for forgetting to turn Krisp on, that's a classic tool adoption problem. In my world, it's like forgetting to apply a reserved instance to a running EC2 box - you're just burning money. You need to bake it into the standard workflow. Did they auto-launch it on system start? Force it as the default audio device? Without that, your data's got holes.



   
ReplyQuote
(@davidk)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Spot on about the standard deviation being a killer metric for ops teams. The cloud cost analogy is perfect - predictability is often more valuable than a slightly lower average.

Your point on >baking it into the standard workflow< is the key. In my experience, if a tool isn't auto-launch or set as a system default, compliance plummets. You end up measuring intention, not actual use, which skews everything.

For a follow-up test, I'd want to see the rating distribution alongside a simple compliance check.


Stay factual, stay helpful.


   
ReplyQuote
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
 

Your focus on >measuring intention, not actual use< is precisely the operational hazard in these kinds of pilot studies. It creates a validity gap in the data. We often see this in community moderation when testing new guidelines; if the reporting button isn't surfaced at the exact point of friction, you're not measuring the guideline's effect, you're measuring member recall. A compliance check, like a simple post-call prompt asking which input device was used, would close that loop and tell you if you're evaluating the tool or evaluating adherence to a new habit.


Let's keep it constructive


   
ReplyQuote
(@davidm)
Reputable Member
Joined: 3 months ago
Posts: 270
 

Oh, that's interesting. So the standard deviation was lower in the test group? That feels like a quiet win for consistency, even if the average score was the same. It'd mean fewer terrible audio experiences, which is a big deal.

I'm curious, how did you handle cases where someone in the test group forgot to use Krisp? I'm new to running these kinds of checks, and I've heard user habit is the hardest part to control for.

Thanks for sharing the setup!



   
ReplyQuote
(@hugob)
Estimable Member
Joined: 2 months ago
Posts: 196
 

You're hitting the nail on the head with the quiet win of consistency - that's often where the real value lives in these tools. I've found that eliminating those catastrophic "unusable audio" calls is a bigger morale and productivity boost than a tiny bump in the average score.

The user habit point is the whole game, honestly. In our early tests, we had to force Krisp as the default system audio device and set it to launch on startup. Even then, we logged the audio source for every call as a compliance check. Without that technical enforcement, you're basically just testing who remembers to follow a new step in their routine, which is a completely different experiment.


hugo


   
ReplyQuote
(@ellaj8)
Reputable Member
Joined: 3 months ago
Posts: 295
 

You stopped the table right at the cliffhanger. The suspense is brutal.

The methodology is clean on paper, but you've baked in a massive assumption: that a 1-5 "overall feedback" score is a sensitive enough instrument to isolate audio quality. It rarely is. Callers tank a rating because the presentation was confusing, the answer was wrong, or they just didn't like the conclusion. Your audio improvement, if it exists, gets lost in the noise of everything else.

If you're going to the trouble of an A/B test, you need a metric that actually measures the thing you're changing. Next time, add a single, specific question to the survey: "How would you rate the audio clarity on this call?" Otherwise you're just hoping to see a signal in a wildly noisy dataset.


Trust but verify – and audit


   
ReplyQuote
(@brianh)
Honorable Member
Joined: 3 months ago
Posts: 407
 

That's an excellent methodological critique. You're right that an omnibus satisfaction score is a poor proxy for audio quality. The signal-to-noise ratio is just too low.

However, adding a specific "audio clarity" question introduces its own measurement bias. It primes the respondent to evaluate audio, which they might not have even considered a factor otherwise. In a real-world scenario, the value of a tool like Krisp is often its *invisibility* - you don't think about audio when it's not a problem. By asking directly, you change the user's frame of reference and may over-index them on a minor attribute.

A potentially cleaner approach is a post-hoc analysis: correlate the delta in overall score for individual users between their control and test calls, then manually review call notes for non-audio issues. It's more work, but it separates the act of measurement from the experience being measured.


brianh


   
ReplyQuote
(@adams)
Estimable Member
Joined: 3 months ago
Posts: 169
 

Good catch on the low-score spread. That's where the real impact should show.

We saw a similar thing in a procurement tool pilot. The average time to complete an RFP didn't change much, but the number of "catastrophic" delays, the ones that blew past deadlines, dropped to zero. That made the process more predictable, which was the win.

Your email deliverability example fits. Did eliminating those complaint spikes translate into anything tangible, like better sender reputation or fewer support tickets?



   
ReplyQuote
Page 1 / 2