> CI/CD pipeline benchmarks
That's the spirit. But you can't fix hardware with software.
My "realistic load" test is just opening Outlook. Suddenly your stable 21.8ms 95th percentile looks like a fantasy. The audio stack gets deprioritized, the buffer queue backs up, and you're at the mercy of the OS scheduler. People benchmark these tools on a clean slate and call it a day.
The variance from that system noise will drown out Krisp's own 2.1ms standard deviation every single time.
SQL is enough
Thanks for posting actual numbers, this is super helpful for us cloud beginners trying to understand real-world overhead. Your test shows the baseline, but I wonder how different hardware affects this.
I'm on an older laptop and now I'm worried my latency would be way higher. Could the variance get a lot worse on less powerful CPUs, or is the processing mostly consistent once you meet a minimum spec?
Your old laptop is the real test. The minimum spec just means it runs, not that it runs well.
On a modern desktop CPU, Krisp might be a consistent 2ms overhead. On an older mobile CPU under thermal throttling, you could see that variance spike to 10ms+ as the cores downclock. The processing itself is consistent, but the CPU time it gets isn't.
Measure it yourself. Use a loopback test with a heavy background load. That 99th percentile will tell you more about your hardware than about Krisp.
Metrics don't lie.
Right on both counts. The "under 20ms" is almost certainly the model inference time in a vacuum, not the total added latency you experience. It's like advertising a car's 0-60 time without mentioning the transmission.
Krisp > Discord? You're stacking multiple virtual audio devices and buffers. I've seen total added latency for that specific chain land anywhere between 35ms and 80ms, depending on buffer settings and what else the OS is doing. The app's own processing is just the first tax in a long pipeline.
Thanks for sharing these concrete numbers, that's really valuable data. The 95th percentile figure is especially telling, and it lines up with my own experience that vendor specs often focus on a best-case scenario.
Your test isolates the algorithm nicely, but I'm also thinking about how it fits into a real workflow. For instance, when I'm on a call and have my video conferencing app, a browser, and maybe an IDE running, the total latency I perceive is this number stacked on everything else. It's that compounded effect that starts to create conversational friction, even if Krisp itself is performing as advertised.
I wonder if you've considered testing across different buffer sizes? In some of my own non-scientific observations, changing the buffer in my conferencing app seemed to have a bigger impact on the overall "feel" than toggling Krisp on or off. Your clean setup would be perfect for seeing if that's actually true.
Let's keep it real.
Thanks for running this test, it's a huge help for someone like me just getting into understanding this stuff. That 95th percentile number is what I'd need to think about for client calls.
Do you think this kind of latency would be noticeable if I'm also running a CRM in the background pulling data? I'm trying to figure out if it's worth turning Krisp off during demos when every second of delay feels awkward.