The overhead for managing remote recorders is real, but it's a known quantity. You can set a noise floor threshold and reject submissions that don't meet it. The inconsistent acoustic distortion from bad home setups is still more coherent and learnable than the perfect sterility of synthetic audio.
You can't spec for the subtle, correlated artifacts a real voice creates when it interacts with a real mic. You either have them or you don't. Synthetic data lacks that causal relationship, and layering noise on top can't recreate it.
—AF
Your results on consistency and acoustic distortion match our findings exactly. That perfect consistency is a training liability. We had to inject random small perturbations into the synthetic audio's timbre and timing before adding any background noise, otherwise the model would overfit to an acoustically sterile speaker that doesn't exist in the wild.
For the config you're using, try adding a jitter parameter to the pitch and timing, even if it's just +/- 2%. It breaks up that machine-precise rhythm and makes the subsequent noise augmentation more effective.
IntegrationWizard
The jitter parameter is a smart adaptation. We implemented something similar after our initial synthetic batches performed poorly in validation. However, we found the order of operations mattered greatly.
If you apply the jitter after you've layered in your environmental noise, the effect is minimal. The model still latches onto the underlying perfect spectral envelope. You have to introduce timing and pitch variance at the waveform generation level, or as a very first preprocessing step, to truly break that coherence. It adds another layer of complexity to the pipeline, but it's necessary.
This does bring up a contract question for API users, though. If you're using a service like Resemble, do their bulk generation endpoints even allow you to specify these low-level jitter parameters per utterance, or are you stuck with the consistency of their core model and forced to post-process? That detail can determine if this approach is feasible.
RTFM — then ask for the audit
Your results mirror what we've seen in production deployments, particularly that critical gap between clean-room WER and real-world generalization. The 8% improvement evaporating in noisy conditions is a classic synthetic data pitfall.
Regarding your config, the lack of acoustic variation is the core limitation. Even with post-hoc noise layering, the underlying spectral coherence of the synthetic voice remains, creating an uncorrelated artifact that models learn to ignore. Your note about needing to add steps for reverberation and background noise highlights the compounding cost problem: you're paying a premium for pristine audio only to spend more engineering time degrading it in a physically unrealistic way.
For command phrases specifically, have you benchmarked against simply using a vocoder-based perturbation pipeline on your original two hours of real data? Techniques like cycle-consistent GANs for voice conversion can sometimes yield more acoustically coherent variations than full TTS synthesis, at a fraction of the API cost, though they introduce their own training complexity.
Mike
Your config block got cut off, which is the only part I actually care about. If you're not sharing the exact voice parameters and which API tier you used, we're just armchair quarterbacking the cost.
That 8% WER drop on the clean set is the siren song of these services. The moment you try to make it useful by adding realistic noise, you're building an entire post-processing pipeline to degrade a premium product. It's architecturally absurd. You'll spend more engineering hours simulating a bad microphone than you would just collecting real data in imperfect conditions.
For command phrases, you're better off taking that small real dataset and running it through a chain of aggressive, structured perturbations:
- Speed/pitch shifts within a narrow band (0.9x-1.1x)
- Convolution with impulse responses from actual rooms
- Non-stationary noise injection where the noise floor correlates slightly with voice amplitude
That'll cost you pennies in compute and often generalizes better than sterile synthetic audio ever will.