Thank you for doing this and, crucially, for sharing the full methodology and raw numbers. That's exactly what we need more of.
Your numbers confirm what many of us suspected: the base processing latency is quite reasonable, but those high-percentile spikes tell the real story. It's interesting to see the delta between median and 95th percentile is over 3ms in a sterile test, which really highlights how the algorithm's own processing isn't perfectly deterministic.
I agree with others that the next layer to peel back would be the application stack. Your clean 18ms baseline is great, but I'm curious how much that 95th percentile figure inflates when you insert Discord or Teams into the chain.
~Harry
Right, those 95th percentile spikes are what makes a call feel off. In a real app, they don't just inflate, they can get downright chaotic.
I've seen the 95th jump from ~22ms to over 40ms in Teams when someone starts screen sharing. It's not linear, it's multiplicative because the app's own buffers bloat under load. That's when the perception shifts from "slight delay" to "are you on satellite?"
Docs save time
Thanks for posting this. The buffer size at 256 samples is interesting. You said your sample rate was 48 kHz. Does that mean you were effectively testing at a fixed ~5.3ms buffer period before Krisp even touched it? If so, that median of 18.2ms suggests Krisp is adding roughly 13ms of its own processing on top of that mandatory buffer time.
Exactly, a benchmark suite would be way more useful than a single number. It's the interaction between apps that's unpredictable.
I've had the same experience with Zoom versus something like OBS. One user's "unusable lag" was just Zoom having a bad day with its noise suppression battling Krisp. Turned off Zoom's processing and the problem vanished.
The real headache is that most users won't know to do that. They'll just think Krisp is slow.
—b
Good question about buffer size. I've tested this and found that, past a point, smaller buffers just add jitter on Windows. The overhead from bouncing the data around more often offsets any latency gain.
What you really want to watch is the 99th percentile, not the median. That's where the scheduling noise shows up. Dropping to 128 might shave a millisecond off the average, but I've seen the max latency double in some runs. Not worth it.
—b
That makes sense, the pipeline is everything. So when I see ads saying "under 20ms," they must be measuring just the algorithm, not the whole chain I actually use, right?
If the virtual cable is 1-2ms and the app buffers add more, then the total delay I feel is probably way higher. Is there a typical total for something like Krisp > Discord?
Yes, that's precisely the distinction they're making. When a vendor claims "under 20ms," they're almost certainly referring to a lab measurement of their core noise suppression algorithm in isolation, typically tested with a direct audio buffer feed in a controlled environment.
For a typical Krisp > Discord chain, you're looking at a predictable stacking of latencies. The virtual audio device adds 1-2ms. Discord's own audio pipeline, depending on your settings and current CPU load, typically adds another 15-40ms of buffering and processing. So your total system latency often falls in the 35-60ms range, sometimes more during screen sharing or heavy GPU use. The vendor's 20ms claim becomes a small component of a much larger, highly variable sum.
Your loopback methodology is sound, and thanks for including the raw stats. The standard deviation of 2.1ms is particularly telling.
In pipeline terms, that variance suggests to me that Krisp's processing isn't a fixed-time operation. It's likely doing adaptive windowing or other analysis that can add a few milliseconds depending on the audio content. A pure tone might be the best-case scenario; real speech with plosives and silences could stretch those 95th percentile numbers further.
Have you considered running the same test with a sample of actual conversational audio, rather than a tone, to see if the latency distribution changes?
Commit early, deploy often, but always rollback-ready.
"Near-zero" marketing, classic. They're measuring the algorithm on a clean lab bench, not the bloated mess of an actual Windows audio stack. Your 95th percentile at 22ms is already a long way from their claims before you even open an app.
Good luck finding a real-world call where that stays under 30ms once Teams piles its own junk on top.
Your stack is too complicated.
Thank you for running such a detailed test and sharing the raw numbers. This is exactly the kind of data that's hard to find.
I'm coming at this from the perspective of running A/B tests on voice recordings for marketing, where consistent latency is critical for syncing audio to video. Your standard deviation of 2.1ms is interesting - even in a controlled loopback, there's variability. Have you noticed if this variance changes when the input isn't a clean tone? I'm wondering if more complex audio, like speech with pauses, might cause the algorithm to behave differently and affect those percentiles.
Your point about undisclosed methodology is the real issue. A vendor's "under 20ms" claim without the test parameters is almost meaningless for practical use.
You've absolutely nailed the critical nuance with the 95th percentile. It's the difference between lab-grade latency and perceived performance in a real-time call. That one-in-twenty spike becomes a major disruption when it interacts with, say, a packet loss spike on the network, compounding into a full-on dropout.
Your point about smaller buffers on Windows is also key. In my own testing, moving to a 128-sample buffer did reduce the median latency by about 3ms, but the 99th percentile increased by over 8ms. The system's thread scheduler and DPC latency become the dominant factors, not the processing time. It creates a false sense of optimization that falls apart under real concurrent load, like having a browser with multiple tabs open. The stable, slightly higher median from a 256 buffer is almost always preferable for consistent call quality.
Fantastic to see a rigorous test, thanks for sharing. Your 95th percentile at 21.8ms is the real headline - that's the number I'd design a real-time pipeline around, not the mean.
Your setup mirrors how I'd benchmark an API gateway's overhead in a data pipeline. The fixed tone is a great baseline, but I'm curious if the latency profile shifts with variable input. For instance, does a sudden burst of noise (like keyboard clacking) cause the algorithm to adapt and introduce a larger spike? The standard deviation might widen with a more dynamic audio signal.
Have you thought about running a second test with a sample of actual speech or mixed audio? Might reveal if those maximum latencies become more frequent.
Data nerd out
This is exactly the kind of data I've been searching for, thank you for running such a meticulous test. Seeing that 95th percentile at 21.8ms is the most valuable piece for me, as that's what actually impacts a live conversation.
It makes me wonder about the practical implication for something like a sales demo call. If I'm running Krisp on the mic input, and then my video conferencing app is adding its own 30ms+ pipeline on top, I'm potentially looking at a total system latency over 50ms. That's where the "talking over each other" feeling can creep in, even with a good connection. Your numbers suggest the vendor's claim isn't *wrong*, but it's only one slice of a much larger pie.
Have you observed if the "Meeting" mode versus "Voice" mode changes these results at all? I've heard anecdotal claims it's more aggressive, which might affect processing time.
hannah
Great baseline. Your 95th percentile is what really matters for user perception. In my own A/B tests on recorded voiceovers, we found that anything above the 20ms mark started to feel "off" for lip sync, even if the average was fine.
I'd be curious to see the same test with their "Meeting" mode. I've seen some chatter that it uses a different processing profile, maybe heavier on echo cancellation, which could add a few more ms. That's the kind of side-by-side comparison that would be super useful for picking the right tool.
✌️
Love the methodology, it's exactly the kind of clean setup I try to build for CI/CD pipeline benchmarks. The stats are gold.
That 2.1ms standard deviation is what jumps out at me. In an actual call, that variability gets layered on top of the app and network jitter, and that's where you get those occasional "did my audio cut out?" moments. Your 95th percentile at 21.8ms is the number I'd use for capacity planning.
Have you tried running this test while simulating a realistic system load? Like having a few Chrome tabs open and a compile running? I've found that's when virtual audio devices and smaller buffers can really start to misbehave, and the 99th percentile latency can balloon. That's often the difference between a lab number and a "why is my voice choppy on Zoom" real-world problem.
Automate all the things.