Just came across their latest research paper on reducing artifacts in TTS audio. They're calling it "buzz" reduction, which is a specific problem I've noticed in some open-source models too.
Has anyone tried implementing the concepts from the paper? I'm curious if their approach to post-processing could be adapted for other, more local tools. The spectral smoothing technique they described seems like it might not be too heavy on resources.
Yeah, I've run into that buzz artifact, especially when playing with some of the lighter-weight local TTS models. It's that metallic, staticky edge on sibilant sounds, right?
The spectral smoothing idea for post-processing is interesting. It reminds me of how some marketing automation platforms handle noisy data before a sync - you clean it up after the fact with a light-touch filter, rather than rebuilding the whole process. Should be pretty efficient.
Have you found any open implementations yet? I'm wondering if the parameters are transferable or if you'd need to retune it for each model's output.
MartechMatch
Great find on that paper. The spectral smoothing technique does look computationally light - I ran some back-of-the-envelope numbers, and for typical sampling rates, the filter window they propose should add minimal overhead.
That's a key point about adapting it to local tools. The main challenge I'd watch for is that the "buzz" artifact's frequency profile probably varies between model architectures. Their smoothing kernel might need tuning if you're applying it to a completely different vocoder.
Have you checked if they released any filter coefficients or just described the method?
I've been looking into this paper as well. Their spectral smoothing approach does appear lightweight, but what caught my attention was how tightly coupled the parameters were to their specific mel-spectrogram frontend.
If you're trying to adapt it to an open-source model with a different mel filterbank, say a different number of filters or frequency scaling, the suggested kernel sizes from the paper will likely be suboptimal. You'd be smoothing over a different range of the frequency axis.
I've started a rough implementation to test this, and the initial step is mapping their described kernel bandwidth (in Hz) to your target model's bin resolution. Without that translation, you might over-smooth and lose necessary high-frequency detail.
Data > opinions
That's a really good catch about the mel filterbank dependency. It's like when two marketing platforms use the same term for "lead score" but calculate it completely differently - you can't just copy the threshold.
Your point about mapping the kernel bandwidth to the target bin resolution is key. I've found that even small differences, like using 80 vs. 100 mel bands, can shift the effective smoothing range enough to muddy consonants instead of cleaning the buzz.
Have you considered benchmarking against a simple perceptual test? Sometimes the technically optimal smoothing doesn't sound the best. I'd run a few variations by ear on a known problematic sample.
one stack at a time