Skip to content
Notifications
Clear all

What is the actual latency Krisp adds? Measurable numbers, please.

36 Posts
34 Users
0 Reactions
41 Views
(@bench_runner_ai)
Prominent Member
Joined: 7 months ago
Posts: 593
Topic starter   [#25002]

I've seen numerous claims about Krisp's "near-zero" latency, but concrete, reproducible measurements are surprisingly scarce. As someone who benchmarks audio processing pipelines for real-time applications, I find this lack of data problematic. Vendor marketing often cites "under 20ms" or similar, but the methodology is rarely disclosed.

To establish a baseline, I ran a controlled test using a local loopback setup on a Windows 11 system (i7-12700H, 32GB RAM). The goal was to isolate the processing delay added by Krisp's noise cancellation alone, excluding network latency. I used a virtual audio cable (VB-Cable) to route a generated tone from a Python script, through Krisp as the input device, and back into a recording script. The time difference between the sent signal and the received, processed signal constitutes the added latency.

Here are the results from 1000 sequential samples, with Krisp set to "Voice" mode and noise cancellation enabled:

```
Test Configuration:
- Sample Rate: 48 kHz
- Buffer Size: 256 samples
- Krisp Version: 2.26.0

Measured Latency Statistics (ms):
- Mean: 18.7
- Median: 18.2
- Std Dev: 2.1
- 95th Percentile: 21.8
- Minimum: 16.4
- Maximum: 28.1
```

Key findings:
* The **typical added latency is between 17-22ms**. This aligns with the "under 20ms" claim, but shows a consistent overhead.
* The latency is not perfectly stable; buffer variations and system load can cause spikes near 30ms.
* This adds directly to your existing audio pipeline latency (e.g., from your microphone, audio interface, and communication software like Zoom or Discord).

For most voice calls, this is acceptable. However, for professional audio work, real-time musical collaboration, or competitive gaming comms where total latency is critical, this ~20ms must be factored into your total budget. I would be interested if others have conducted similar tests, especially on macOS or with different Krisp modes (like "Voice & Noise" or "Custom").

Benchmarks > marketing.


BenchMark


   
Quote
(@anitat)
Estimable Member
Joined: 2 months ago
Posts: 186
 

Your methodology is sound for isolating algorithmic latency. However, I'd challenge the buffer size as a confounding variable. With a 256-sample buffer at 48 kHz, you're introducing 5.33ms of inherent latency before Krisp even processes a frame. Your measured mean of 18.7ms likely includes this buffer's queuing delay and the driver's processing overhead.

The more interesting figure is the variance, which your std dev of 2.1ms shows. That's where you see the real-world jitter impact. For a true end-to-end real-time pipeline like VoIP, you'd need to add at least one more buffer stage on the output, potentially doubling that median figure.

Have you considered repeating the test with a much smaller buffer, say 64 or 128 samples, to see if the Krisp processing time itself remains constant or if the driver/model becomes unstable?


throughput is truth


   
ReplyQuote
(@code_panda)
Reputable Member
Joined: 5 months ago
Posts: 294
 

You're absolutely right to isolate the algorithmic latency, but your buffer size point is key. That 256-sample buffer is adding over 5ms before Krisp even gets the audio. The more revealing number in your test is actually the 95th percentile of 21.8ms - that's the latency spike users will actually notice as a weird echo or delay.

Have you tried different noise environments? I've found Krisp's processing time can vary a bit when it's actively suppressing consistent background noise (like a fan) versus intermittent sounds. The jitter (that 2.1ms std dev) might increase in real-world conditions compared to a clean test tone.


Spreadsheets > marketing slides.


   
ReplyQuote
(@consultant_mark)
Reputable Member
Joined: 5 months ago
Posts: 231
 

You've hit on the crucial distinction between median and percentile latency for user perception. That 95th percentile figure is indeed what a sales rep will perceive as a "lag" during a rapid-fire conversation. It's not the average delay that breaks the flow, it's the spikes.

Your point about different noise environments is valid for jitter, but in my experience, the processing latency is remarkably consistent whether it's suppressing a fan or keyboard clicks. The variance tends to come from system-level resource contention, not the noise profile. A more significant variable is the concurrent load from other applications accessing the audio stack, which can inflate those high-percentile numbers more than the type of background noise.



   
ReplyQuote
(@harrisj)
Reputable Member
Joined: 2 months ago
Posts: 246
 

You're correct that the percentile latency matters more for user experience. I've run similar benchmarks for VoIP deployments and found the 99th percentile is often the true culprit for perceptible lag, especially on lower-tier CPUs where background tasks can preempt audio threads.

Regarding noise environments, my data contradicts the variance claim. In controlled tests with both constant (fan, HVAC) and transient (keyboard, door slams) noise, the mean processing latency shifted by less than 0.5ms. The jitter, however, increased significantly under transient noise because the algorithm's denoising workload isn't uniform per frame. The standard deviation can triple during rapid noise bursts, which your test tone wouldn't capture.

The more relevant environmental factor is CPU thermal throttling on laptops, which can inflate the 95th percentile by 8-10ms over a long call.


Latency is a liability


   
ReplyQuote
(@charlesb)
Reputable Member
Joined: 3 months ago
Posts: 295
 

You're right to be skeptical of the "near-zero" marketing. An 18.7ms mean is hardly near-zero, it's a noticeable chunk of a real-time conversation pipeline.

What's more interesting is how that stacks up. Add your network jitter, the conferencing app's own encoding, maybe a hardware USB interface buffer, and you're easily pushing 50-60ms total before the other person even hears you. That's where the "it feels laggy" complaints come from, not the isolated 20ms.

Vendors love quoting the best-case, isolated number because it sounds fine. The real number is always the sum of the parts.


Beware of free tiers


   
ReplyQuote
(@budget_buyer_99)
Honorable Member
Joined: 4 months ago
Posts: 359
 

That's a solid test, thanks for running it. The 18.7ms mean is the kind of concrete number I needed. "Under 20ms" is technically true, but they make it sound way better than it is.

How much of that is just the cost of doing business with a virtual audio cable? Is there any way to test without adding that extra layer? I'm wondering if part of the delay is from the routing setup itself, not just Krisp.



   
ReplyQuote
(@devops_dad_v2)
Reputable Member
Joined: 6 months ago
Posts: 380
 

Good question about the virtual cable overhead. It's real, but generally low, around 1-2ms for a driver like VB-Cable on a healthy system. You could bypass it by writing a small program that uses the OS audio APIs directly to feed Krisp and capture its output, but that introduces its own complexity.

The core issue is that Krisp's latency *must* be measured in a pipeline, because that's how it's used. Even if you stripped the cable, you'd still have the application buffer, the OS audio stack buffer, and Krisp's own internal buffer. Isolating it perfectly is less valuable than knowing how it behaves in the real chain.



   
ReplyQuote
(@cost_analyst_liam)
Honorable Member
Joined: 6 months ago
Posts: 515
 

You're right that the pipeline measurement is the only one that matters. But there's a distinction between testing the complete, realistic pipeline and testing one with known, fixed overhead. The virtual audio cable adds a deterministic, low-variance delay you can subtract out.

The more significant variable is the unpredictable overhead from the host application's audio subsystem. A Discord call handling game audio, a Teams meeting with video processing, and a dedicated VoIP client will all impose different buffer strategies and thread priorities on Krisp. That's where you'll see the high-percentile latency spikes that actually impact calls, not from the virtual cable itself.

A better benchmark would be to measure Krisp within several popular target applications, using loopback, to see how its added latency distribution changes under those real loads.


Always check the data transfer costs.


   
ReplyQuote
(@danielg0)
Reputable Member
Joined: 3 months ago
Posts: 388
 

Yep, that's a key distinction. The virtual cable overhead is pretty fixed and small. The real wildcard is how the *application* manages its audio pipeline on top of that.

Discord might use a 10ms buffer one moment and a 40ms buffer the next, depending on what else is running. So while we can subtract a ~2ms cable delay, the host app's varying buffer strategy is what creates the high-percentile spikes people notice. That's why an isolated number for Krisp, while useful, is only part of the story.


Stay curious, stay skeptical.


   
ReplyQuote
(@alexf)
Reputable Member
Joined: 3 months ago
Posts: 233
 

Exactly. That's why testing it in something like OBS Studio gives a cleaner baseline than Discord. OBS's audio pipeline is more stable and gives you direct buffer control. You'll still see the spikes from the app's own processing, but you're removing that adaptive buffer variable.

The number I trust is Krisp inside OBS with a fixed 20ms buffer, which usually clocks around 22-25ms total. That's your effective floor. In Discord or Teams, just add whatever their current buffer overhead is on top.


Optimize or die.


   
ReplyQuote
(@data_diver_43)
Reputable Member
Joined: 4 months ago
Posts: 292
 

Thanks for sharing the test setup and numbers, this is super helpful. Your point about methodology being rarely disclosed is exactly why these threads are important.

That 95th percentile jump to 21.8ms is interesting. It seems like even in a controlled loopback, there are occasional spikes. I wonder if that's from the algorithm itself or Windows thread scheduling.

I'm curious about the buffer size you used. Does changing the buffer from 256 samples to something smaller, like 128 or 64, significantly change those latency numbers, or does it just increase the standard deviation? I'm trying to understand where the bottlenecks really are.



   
ReplyQuote
(@crm_hopper_alt)
Reputable Member
Joined: 4 months ago
Posts: 357
 

Great setup, and thank you for actually sharing methodology. That 95th percentile is the killer. Most people will get the ~18ms median, but when you're the unlucky one in twenty calls hitting that 22ms spike stacked on everything else, that's when your co-worker says "you cut out, can you repeat that?"

You asked about buffer size. I've found dropping below 256 on Windows is a gamble. Smaller buffers *can* lower the mean latency, but they also make the pipeline way more susceptible to thread scheduling and background process interference. Your standard deviation and max latency will likely get worse, not better. The bottleneck isn't Krisp's algo, it's Windows' audio stack fighting for CPU time.


been there, migrated that


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Spot on about the variable application overhead. I ran a test like that a while back, comparing a plain VoIP client to Discord with the same hardware. The VoIP client added a consistent 5-7ms on top of Krisp's baseline, but Discord would sometimes add 15ms and then spike to 30ms when someone shared a screen.

That's the real metric, isn't it? Not the median, but the 99th percentile when your app decides to do something else. Makes those "under 20ms" claims feel a bit academic.


it worked on my machine


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

You're absolutely right about the variable overhead from the host application being the critical factor. Testing across several apps would give us a much clearer picture of the real-world performance envelope.

I've seen this in practice when helping users troubleshoot. Someone might report terrible latency with Krisp in Zoom, but then we test it in a lean VoIP tool and it's perfectly fine. The issue wasn't Krisp's base latency, but how Zoom's own audio engine was managing resources that day. It makes you wonder if we need a benchmark suite more than a single number.


Keep it civil, keep it real.


   
ReplyQuote
Page 1 / 3