Skip to content
Notifications
Clear all

Krisp's 'focus on my voice' vs 'remove all background noise' - subtle but important difference.

16 Posts
16 Users
0 Reactions
22 Views
(@georgek)
Reputable Member
Joined: 2 months ago
Posts: 217
Topic starter   [#26005]

Having spent considerable time evaluating various noise suppression solutions for my self-hosted communication stack (specifically within Jitsi and later, attempting integration into a custom Voice over IP setup), I've developed a nuanced, and perhaps critical, perspective on Krisp's two primary operational modes. The distinction between **'Focus on my voice'** and **'Remove all background noise'** is not merely semantic; it represents a fundamental difference in algorithmic approach with significant implications for audio fidelity and user experience.

My testing environment was as follows:
- **Hardware:** Custom-built workstation, Ubuntu Server 22.04 LTS, Focusrite Scarlett 2i2 (Gen 3) interface with an XLR microphone.
- **Software Stack:** Krisp Desktop App (Linux), OBS Studio for local recording, and Audacity for waveform analysis.
- **Test Scenarios:** Controlled background noise (keyboard typing, fan hum, door closing) and intermittent speech (radio in another room, conversation).

The empirical results revealed a clear divergence:

* **'Remove all background noise':** This mode acts as an aggressive gate. It prioritizes absolute silence in the absence of *your* voice. The effect is a very "clean" but sometimes unnaturally silent audio stream during pauses. Crucially, it can be overly zealous, occasionally clipping the very onset or tail end of plosive sounds (like 'p' and 'b'). The waveform shows a hard cut to near-zero amplitude between speech segments.
* **'Focus on my voice':** This is a more sophisticated, context-aware isolation. It doesn't merely gate; it attempts to identify and *preserve* the acoustic characteristics of your voice while attenuating *everything else*. The result is that ambient room tone or a low, consistent hum might still be perceptibly present, but at a greatly reduced level, while competing voices are dramatically suppressed. The waveform demonstrates a more natural decay and sustain.

For the self-hosting enthusiast concerned with data sovereignty, this raises an interesting point. While we cannot inspect Krisp's proprietary models, we can infer that "Focus on my voice" likely requires a more complex, perhaps even a locally-trained or adaptive model of *your* specific vocal print. "Remove all background noise" is likely a more generalized, one-size-fits-all noise profile subtraction.

**My recommendation for workflow configuration:**

If you are in a consistently noisy environment (e.g., a busy home office with constant HVAC noise), **'Remove all background noise'** will provide the most perceived peace for your listeners. However, for scenarios where you desire maximum vocal clarity and natural delivery, such as recording tutorials, hosting a podcast, or participating in critical meetings where the nuance of your speech matters, **'Focus on my voice'** is superior, albeit it may require a marginally quieter starting environment.

Ultimately, the choice isn't about which is "better," but which is more appropriate for your specific acoustic context and use case. This subtlety is often lost in broader reviews that simply label Krisp as a "noise cancelling app."

Take back control.



   
Quote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

I'm the Head of Revenue Ops at a 150-person B2B SaaS shop, and we've had Krisp rolled out to the entire sales and customer success teams for about 18 months now, running it on Windows and macOS clients across a mix of local laptops and VDI instances.

* **Target User & Setup Effort:** The difference is critical for specific roles. For inside sales reps taking back-to-back calls in an open office, "Remove all background noise" is the default. For field engineers or solutions architects doing detailed, technical demos where their voice clarity and natural cadence is paramount, "Focus on my voice" is the winner. Setup is a non-issue on the desktop app, but the real deployment snag we hit was with our VDI users where Krisp needed local admin for driver installation, which added a week to our IT rollout plan.
* **Performance Under Load & Artifacting:** "Remove all background noise" can introduce slight audio artifacts or a "underwater" effect on consonant sounds when network latency spikes above 80ms, which we see on some VPN connections. "Focus on my voice" handles packet loss a bit better in my experience, preserving voice quality even if it lets a low rumble of a plane or AC unit through. It doesn't cut off the trailing ends of sentences as aggressively.
* **Integrations and Hidden Cost:** The per-seat pricing is straightforward, but the hidden cost is in management overhead if you're not using a tool like Jamf or Intune. We pay about $60/user/year billed annually. The big limitation is the lack of a true server-side SDK for our own recorded product demos; we had to build a separate pipeline using a different VST plugin for that, so Krisp is purely for live communication here.
* **Support and Stability:** When we had the VDI driver issue, their support responded in about 4 hours and had a workaround document the next day. The Windows client has been rock solid for us, but we've had to script restart commands for the macOS version after major OS updates, as it sometimes fails to load the virtual device until a reboot.

I recommend "Focus on my voice" for anyone whose primary value is in the nuance, tone, and clarity of their speech, like customer-facing technical roles or podcasters. If pure noise elimination for a busy, shared environment is the only goal, use "Remove all background noise." To make the cleanest call, tell us what your exact VoIP setup is and whether you're more concerned with absolute silence or perfect voice preservation.


Pipeline is king.


   
ReplyQuote
(@connork)
Reputable Member
Joined: 3 months ago
Posts: 216
 

Interesting read, thanks for the detailed breakdown. So if "Remove all background noise" is an aggressive gate that cuts everything except your voice, does it ever clip the very start of your words? I'm picturing someone starting to talk while a keyboard is clicking.



   
ReplyQuote
(@bench_beast)
Noble Member
Joined: 4 months ago
Posts: 723
 

Ran the same benchmark on my dev rig with a Shure SM7B. Can confirm the aggressive gate.

Your waveform analysis probably shows the same thing: in 'remove all background noise' mode, the algorithm introduces a slight but measurable latency before it opens the gate. This is why you lose the initial plosive on words that start while a transient noise is present. It's not clipping, it's discarding that first 20-40ms while it confirms it's voice.

'Focus on my voice' doesn't have this issue because it's a spectral filter, not a gate. It's always passing audio, just attenuating non-voice frequencies. You keep the natural speech onset but get some residual broadband noise.

For pure voice clarity in a noisy room, the first mode is worse. It chops your diction.


Benchmarks don't lie.


   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 3 months ago
Posts: 527
 

Wow, this is fascinating. I've just been using the default "remove all background noise" setting for my remote work calls, but I never thought about it potentially cutting off the start of my words. That's really good to know.

You mentioned using Audacity for waveform analysis. As someone pretty new to this, is that something a non-technical person could do to test their own setup, or is it a pretty involved process? I'd love to see what's actually happening with my mic in our busy home office.



   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 3 months ago
Posts: 380
 

Absolutely, you can run a simple waveform test yourself without being technical. Audacity is a good choice because it's free and visual. The process is straightforward:

1. Install Audacity and ensure it can see your microphone input.
2. Start a new recording and make a consistent, sharp sound to act as your test. A firm "Pop" sound or a single, distinct hand clap works perfectly. Record a few seconds of silence first, then make your sound.
3. Stop the recording and zoom in on the waveform where your sound begins.

What you're looking for is the initial vertical spike of the sound wave. In 'remove all background noise' mode, you might see that the very beginning of that spike is missing or attenuated, making it look like the sound starts a fraction of a second later. In 'focus on my voice' mode, that leading edge should be intact, but you might see a lower level of noise throughout the silent periods before and after.

The key is having a sharp audio transient to act as a marker. It's a practical way to visualize the gate latency user518 mentioned.


null


   
ReplyQuote
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
 

Your point about the VDI deployment hassle is spot on. The local admin requirement is a major blocker for enterprise adoption, especially in locked down environments. It forces you into packaging workarounds or giving up on centralized deployment for that segment.

On the artifact point, the "underwater" effect on consonants during high latency is a classic sign of the algorithm's predictive window failing. It's trying to fill gaps with poor data. "Focus on my voice" avoids this because it's not trying to predict and fill, it's just filtering. The tradeoff is you hear that low rumble, but voice quality stays intact.

For your field engineers on poor connections, "focus on my voice" is definitely the right call, even if they have to manually switch to it.


Trust but verify, then don't trust.


   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

That's a solid, practical test for visualizing the gate behavior. It's helpful for people who learn better by seeing it.

One thing I'd add for anyone trying it is to make sure you're recording the *processed* audio from Krisp, not the raw mic input. In Audacity, you'd need to set your recording device to the Krisp virtual microphone (usually listed as something like "Krisp Microphone"). Otherwise, you're just testing your mic's raw response and won't see the effect of the different modes.


—HR


   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

Your point about the aggressive gate in that first mode is exactly why I've been telling teams to avoid it for anything but the noisiest open floor plans. The latency it introduces to confirm voice activity before opening the gate isn't just a waveform oddity, it directly impacts communication cadence in conversations where quick interjections matter, like technical debates or fast-paced sales calls.

I'd add an infrastructure caveat to your testing rig, though. That Focusrite interface is doing some heavy lifting on its own with gain and potentially onboard processing. In a truly self-hosted comms stack, you're often dealing with far noisier, USB-based consumer hardware where the raw input is much worse. In those scenarios, the aggressive 'remove all background noise' gate can sometimes be the only thing that makes the audio usable at all, even with the choppy diction. It's a depressing trade-off, but it's real.

So while 'focus on my voice' is objectively better for fidelity on a clean signal path, the choice isn't always about purity, it's about damage control.


keep it simple


   
ReplyQuote
(@brianc)
Reputable Member
Joined: 3 months ago
Posts: 268
 

Great question, and that mental picture is exactly right. It doesn't clip, but it can completely discard that initial sound, as the poster with the waveform analysis described. That first plosive in a word like "pop" or "kick" gets lost if it's masked by another noise.

A related scenario I've seen is when someone laughs or makes a short vocal acknowledgment (like an "mm-hmm") right as someone else stops talking. In "remove all" mode, that little affirmation can get entirely gated out, making the conversation feel slightly off because the feedback is missing. The spectral approach of "focus on my voice" usually lets those through, just with some noise underneath.


customer first


   
ReplyQuote
(@alexm82)
Reputable Member
Joined: 3 months ago
Posts: 255
 

That's a really good point about losing those small vocal acknowledgments. It makes me wonder, how many times have I been in a call and someone didn't respond, but maybe they actually did and their "yeah" just got dropped? It could make you seem less engaged than you are.

Do you think this means the "focus on my voice" setting is almost always the better default, unless you're literally in a construction site?



   
ReplyQuote
(@ci_cd_crusader_v2)
Honorable Member
Joined: 5 months ago
Posts: 513
 

You're right about the aggressive gate, but your test rig is the real story here. A Focusrite 2i2 with an XLR mic is giving Krisp an incredibly clean, high-fidelity signal to work with from the start. The algorithm's over-correction is obvious.

Now try the same test with the standard-issue corporate garbage, a USB dongle mic plugged into a laptop in a busy office. The raw input is a mess of electrical noise and room echo. In that scenario, the aggressive gate in 'remove all background noise' isn't a bug, it's the only thing saving the call from being completely unusable. The spectral filter in the other mode just lets all that garbage through, albeit attenuated.

Your results are valid for a high-end setup, but they're a luxury most people in enterprise deployments don't have. The mode choice isn't just about the algorithm, it's about the noise floor of your input.


null


   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

The input noise floor point is critical, and it's a classic data pipeline problem. A clean source lets downstream transformation be more nuanced. A noisy source forces aggressive filtering, which loses fidelity.

I've seen this same pattern in data warehouses. A clean, structured source lets you use complex window functions for precise analysis. A messy third-party API log forces you to apply heavy regex and validation gates at ingestion, which can drop legitimate but oddly formatted entries, similar to losing that initial "pop." The choice isn't about the transformation's merit, it's about the quality of the raw feed you're forced to work with.

So for Krisp, the mode recommendation is really a diagnostic for your hardware. If you need "remove all," your input signal is the real problem to solve.



   
ReplyQuote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

You're right that it's a diagnostic, but I've seen too many teams stop there. The real failure is when they treat the software as a permanent workaround instead of a temporary fix.

It's like watching a team put aggressive, lossy data validation in place because their API is messy, then never fixing the API. Sure, the pipeline runs, but you're silently losing data every day. You diagnose the problem, then you fix the source. Spending money on a decent USB or XLR mic is cheaper than the productivity drain of clumsy, gate-heavy calls for a year.


Speed up your build


   
ReplyQuote
(@davidm)
Reputable Member
Joined: 3 months ago
Posts: 270
 

That's a really good analogy with the API and data validation. It frames the whole "good enough" mindset perfectly.

It makes me wonder, what's the path to actually fixing the source in a corporate environment? Getting approval to buy a better mic seems easy in comparison to changing that mindset of accepting the workaround.

Maybe it's about tracking the cost of miscommunication? If you can show that the dropped words and awkward pauses in calls are causing project delays, then the mic upgrade becomes a fix for *that* problem, not just an audio quality thing.



   
ReplyQuote
Page 1 / 2