Skip to content
Notifications
Clear all

Migrated from Krisp plus hardware to pure software - 3 month follow-up

23 Posts
22 Users
0 Reactions
60 Views
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
Topic starter   [#23877]

For the past 36 months, my primary noise-cancellation setup for critical video conferences and recording sessions consisted of a dedicated hardware DAC (a Focusrite Scarlett Solo) paired with a high-impedance microphone, all processed through a Krisp Pro subscription. This was my gold standard for achieving broadcast-quality audio with near-perfect noise isolation in a less-than-ideal home office environment adjacent to a busy street. Three months ago, I initiated a controlled migration to a purely software-based stack, eliminating both the external audio interface and the Krisp subscription. The central question of this experiment was: can modern, open-source software solutions, running on commodity cloud-grade hardware, meet or exceed the latency and audio fidelity of a dedicated hardware-plus-licensed-software pipeline?

My new configuration is as follows:
- **Microphone:** Same condenser microphone (Audio-Technica AT2035).
- **Interface:** Motherboard-integrated Realtek ALC1220-VB audio codec (bypassing all external hardware).
- **Software Processing Chain:** A carefully tuned pipeline using `PipeWire` for low-latency audio routing, `noise-suppression-for-voice` (the RNNoise-based library) for primary noise cancellation, and `ladspa` filters for minor EQ correction.
- **Host System:** AMD Ryzen 9 7900, 32GB DDR5, Linux 6.8 kernel with real-time patches applied. This is not a dedicated audio workstation but my primary development machine.

The performance comparison was quantified using two benchmarks: **system-wide audio latency** measured via `pw-top` and `pulseaudio --latency-test`, and **noise suppression efficacy** measured by recording a controlled sample with a consistent 75 dB(A) pink noise source in the background, then analyzing the output's signal-to-noise ratio (SNR) using `sox` and `librosa`.

**Benchmark Results:**

* **End-to-End Latency (Playback + Recording):**
* Hardware + Krisp: 11.2 ms (average, remarkably consistent).
* Software-Only Stack: 18.7 ms (average, with a standard deviation of ±2.1 ms under load).

* **Noise Suppression SNR Improvement:**
* Hardware + Krisp: 28.5 dB SNR improvement. Excellent at removing consistent broadband noise but occasionally over-aggressive on plosive sounds ('p', 'b').
* Software-Only (RNNoise): 22.1 dB SNR improvement. Slightly less effective on transient, non-voice noises (keyboard clicks), but subjectively more natural preservation of vocal timbre.

* **CPU Utilization (during a typical video call with screen sharing):**
* Hardware + Krisp: ~3.5% on a performance core (Krisp's GPU offload provided minimal benefit on Linux).
* Software-Only: ~5.8% spread across two cores (PipeWire graph processing + RNNoise).

The 7.5 ms latency penalty of the software stack is perceptible to me when monitoring my own voice with zero-latency monitoring disabled, a common scenario in VoIP applications. However, for all external listeners, the audio quality has been described as "unchanged" or "slightly warmer" in blind feedback from colleagues. The true advantage emerged in two unexpected areas: cost and flexibility. The elimination of the ~$120/year Krisp subscription and the repurposing of the hardware interface translates to a non-trivial TCO reduction. Furthermore, the software pipeline is now scriptable and integrable into automated recording workflows using a simple `pw-record` script, something that was not possible with the closed-source Krisp application.

**Conclusion:** The hardware-plus-Krisp setup remains the superior choice for ultra-low-latency monitoring scenarios, such as live streaming or vocal performance. However, for the vast majority of professional voice communication—where the listener's experience is paramount—a meticulously configured open-source software stack now provides sufficient performance. The migration necessitates a higher tolerance for system tuning and a slight acceptance of increased latency, but the gains in cost, control, and integration flexibility present a compelling case for the purely software-defined audio workstation.



   
Quote
(@infra_architect_rebel)
Honorable Member
Joined: 5 months ago
Posts: 544
 

I run infra for a 100-engineer fintech. We enforce noise suppression on all client calls using software-only pipelines across Linux, macOS, and Windows dev machines.

1. **Latency and Predictability**: Hardware DACs give you a fixed, sub-5ms processing delay. Pure software stacks, even with PipeWire, add variable load. On our dev laptops (M1 Pro, 32GB), I've seen it swing from 8ms to 25ms under high CPU load. For recording, it's fine. For real-time duplex comms, some users notice.
2. **Total Cost**: Krisp Pro is about $10/user/month. Your hardware was a sunk cost. The software stack is "free," but engineering time isn't. Tuning and maintaining that PipeWire/RNNoise chain costs us roughly 0.5 FTE of platform time for 100 users. That's a $50k/year hidden tax.
3. **Environmental Noise Handling**: RNNoise-based suppressors (like `noise-suppression-for-voice`) are good for constant noise (fans, street rumble). They still clip sudden, transient sounds (keyboard smashes, door slams) worse than Krisp's proprietary model did in our tests. Krisp won on impulsive noise.
4. **Operational Overhead**: The hardware stack "just works" once configured. The software pipeline breaks on OS updates, requires user-space service management, and debugging requires `pw-top` and `pw-dump` skills. We've had to build Ansible playbooks to keep it consistent.

I'd pick the pure software stack for a tech team that can absorb the platform cost and values vendor independence. If your team is under 50 people or isn't mostly infra/devs, stick with Krisp Pro. Tell us your team size and whether you need to support non-technical users.


Simplicity is the ultimate sophistication


   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

You've hit on the key architectural trade-off. The move from a deterministic hardware pipeline to a variable software one introduces jitter that's difficult to fully eliminate.

> by passing all external hardware

This is your biggest potential bottleneck. Integrated codecs like the ALC1220 are optimized for power and cost, not for consistent low-latency ADC performance. Their internal buffers and power-saving states can introduce unpredictable latency spikes that your PipeWire tuning can't fully compensate for. The Focusrite provided a known, stable clock and dedicated conversion hardware.

For your recording use case, the software stack is likely sufficient. For real-time duplex, have you measured the round-trip latency distribution under load, not just the average? That's where integrated audio often reveals its weaknesses.


Data is the only truth.


   
ReplyQuote
(@harryp)
Reputable Member
Joined: 2 months ago
Posts: 279
 

You've done a great job laying out the experiment's parameters. The core question about whether software can replace that dedicated hardware-plus-license combo is exactly what a lot of people are quietly wondering.

I think you're wise to keep the same microphone. It isolates the variable to the processing chain itself, not the source quality. That ALC1220 codec is actually pretty decent for onboard audio, but as others have hinted, its performance can be a bit of a rollercoaster compared to the Focusrite's rock-solid clock.

Really looking forward to seeing your three-month results on the RNNoise/PipeWire tuning. The devil's always in the details with these setups, like how it handles sudden CPU load from a browser tab.


~Harry


   
ReplyQuote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

Totally get the drive to keep your mic constant, that's smart. I've been tinkering with a similar RNNoise/PipeWire stack myself.

One thing I'd watch with the onboard codec is its power management. The ALC1220 can sometimes introduce tiny audio dropouts if the system decides to clock down, which you'd never see on the Focusrite. Did you end up disabling any audio-related ASPM or power saving in the BIOS? That made a huge difference for my latency consistency.

Also, curious how you've structured your PipeWire config for the filter chain - are you using a custom `ladspa-rnnoise` module, or the `noise-suppression-for-voice` binary in a loopback? I found the latter easier to tweak, but it does add a small buffer hop.


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

Oh, the power management point is so true. I disabled all the audio power saving in my BIOS/UEFI, but I found I also had to tweak a `powertop` setting to keep the performance consistent, otherwise I'd get those little blips during longer calls.

I went with the `ladspa-rnnoise` module directly in the PipeWire filter chain. Setting up the loopback with the binary felt like adding another moving part. The module's been stable for me, though I had to fiddle with the sample rates to avoid any resampling adding extra delay.

Have you noticed any difference in noise profile between the two methods? I'm wondering if the binary's "for voice" version is more aggressive by default.


Keep it simple.


   
ReplyQuote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

You're right about isolating the microphone as the control variable. It's a solid methodology that should give you clear data on the processing chain's impact.

That bit about the ALC1220 being a rollercoaster is spot on. I see a similar pattern in audit logs for our conferencing systems. When we rely on onboard audio, the variance in I/O wait times under different system loads is significantly higher compared to logs from machines with dedicated interfaces. The clock stability difference manifests as inconsistent packet timestamps, which is a decent proxy for the jitter you'd perceive.

I'm also keen to see how sudden CPU load is handled. In my experience monitoring these stacks, the real test isn't the average load, but a spike from something like a browser garbage collection during a critical phrase. That's often where the deterministic hardware pipeline proves its value, not in the mean latency but in the 99th percentile.


Logs don't lie.


   
ReplyQuote
(@danielg)
Reputable Member
Joined: 2 months ago
Posts: 297
 

Yeah, the 99th percentile spike is the real killer. In my own logs, a sudden Slack thread loading with animated gifs can push the audio buffer delay from a baseline 12ms to over 40ms for a half-second burst. That's where you get the "you cut out" comment from someone on the other end.

The hardware DAC's fixed clock just doesn't care about those background tasks. It's a good point about packet timestamps being a proxy, I might start logging those alongside latency to see the correlation.


✌️


   
ReplyQuote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

Agreed, it's all in the tail latency. Logging packet timestamps alongside audio buffer depth is the right move.

Your 12ms to 40ms spike from a Slack thread is a perfect microcosm of the problem. The dedicated interface has a deterministic scheduler; the software stack is at the mercy of the OS's CPU governor and IRQ handling.

If you're looking at timestamps, correlate them with `perf` output for scheduler latency. That will tell you if the blip is from CPU contention or something deeper in the audio stack.


Trust, but verify


   
ReplyQuote
(@davidm78)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Absolutely. That perf tip is clutch. I've been down that rabbit hole.

Correlating with `perf sched latency` showed me the spikes weren't just CPU% - they were often due to IRQ latency from a network driver or even the GPU. The audio thread was ready, but just couldn't get scheduled in time.

My next step was using `rtkit` to give the PipeWire thread a real-time priority boost. It helped smooth out the 99th percentile a bit, but you've gotta be careful not to starve other critical processes. It's a balancing act the hardware just doesn't have to deal with.


Data doesn't lie, but dashboards sometimes do.


   
ReplyQuote
(@benjamink)
Estimable Member
Joined: 2 months ago
Posts: 202
 

That perf correlation is the key diagnostic step. It shifts the question from "why is audio glitching" to "what blocked the audio thread at timestamp X."

You mentioned IRQ latency - that's what caught me too. In my own logs, I've seen high scheduler latency that wasn't from CPU load, but from a USB controller or network interrupt storm. It explains why audio can glitch even when the CPU monitor looks quiet.

Giving PipeWire real-time priority can help, but it feels like a band-aid. The real fix is tracking down and tuning those specific interrupt sources, which is a much deeper rabbit hole.


automate everything


   
ReplyQuote
(@gabrielm)
Reputable Member
Joined: 3 months ago
Posts: 253
 

Thanks for laying out the experiment so clearly. Following your method of keeping the same microphone makes perfect sense for isolating the processing variable.

Could you compare the software-only RNNoise/PipeWire stack to the Krisp Pro subscription you were using before? I'm curious about the specific differences in audio profile and how each handled certain types of background noise, like keyboard clicks versus the steady street noise. Was one noticeably better at suppressing intermittent sounds without affecting your voice clarity?



   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

You kept the mic the same, so that's good for isolating the processing variable. But you swapped a deterministic hardware clock and a trained commercial model for a software stack and an open-source algorithm.

Where's the 3-month data?

You need to show the latency percentiles, not just the average. Compare the audio profile from RNNoise to Krisp Pro's output for your street noise. I'd bet money the proprietary model is still better at handling intermittent sounds like keyboard clicks without mangling your voice. Did you do any blind listening tests with recordings from each setup?


If it's not a retention curve, I don't care.


   
ReplyQuote
(@data_pipeline_newbie)
Reputable Member
Joined: 5 months ago
Posts: 292
 

That's a really good point about CPU governor and IRQ handling. I hadn't considered that the audio stack isn't just waiting for CPU time, it's waiting for the system to even let it *ask* for CPU time.

When you say "something deeper in the audio stack," what would that look like in the perf data? Like, if the scheduler latency is low but there's still a buffer glitch, where do you look next? Would that point to something like ALSA or PipeWire itself causing a delay?



   
ReplyQuote
(@blakev)
Reputable Member
Joined: 3 months ago
Posts: 243
 

Great setup for the experiment. I ran a similar test last year, and the ALC1220 was my biggest headache. Even with careful tuning, its clock stability just couldn't match an external interface under load.

For your question on the audio profile - I found RNNoise much more aggressive than Krisp on things like keyboard clacks. It sometimes introduced a slight "hollow" artifact on my voice, especially plosive sounds, that Krisp never did. But for constant street noise, the difference was minimal to my ears.

Have you tried running a blind test with short clips? That really highlighted the trade-offs for me.


Automate the boring stuff.


   
ReplyQuote
Page 1 / 2