Skip to content
Notifications
Clear all

Help: Resemble's API keeps rejecting my audio clips as 'low quality' - what's the actual spec?

21 Posts
19 Users
0 Reactions
49 Views
(@chrisr)
Reputable Member
Joined: 3 months ago
Posts: 227
Topic starter   [#27309]

I've been working on a project to integrate Resemble AI's voice cloning and generation capabilities into a customer service automation platform. The goal is to create dynamic, natural-sounding responses. However, I've hit a consistent roadblock during the voice cloning phase: approximately 70% of my submitted audio clips are rejected with a generic 'low quality' error from their `/api/v1/clone` endpoint.

The documentation states clips should be "high quality" and provides basic parameters like WAV format and a minimum length, but the rejection feedback is non-specific. Given my background in observability, I'm accustomed to working with precise, measurable thresholds. The lack of concrete, technical specifications is making this process iterative to the point of inefficiency.

I've conducted a series of controlled tests using `ffprobe` and `sox` to analyze my submissions, comparing the ones that passed versus those that failed. Here's a summary of my observations on the clips that *were* accepted:

* **Format:** Linear PCM WAV, as stated.
* **Sample Rate:** 22050 Hz appears to be the sweet spot. I had one clip at 44.1 kHz rejected, while an identical transcript resampled to 22050 Hz passed.
* **Bit Depth:** 16-bit.
* **Channels:** Mono (1 channel). Any stereo file I submitted was rejected.
* **Background Noise:** A quiet, consistent noise floor (~-60 dBFS) was tolerated, but any clips with dynamic noise (keyboard clicks, variable fan hum) were rejected. This suggests a non-standardized SNR (Signal-to-Noise Ratio) requirement they aren't publishing.
* **Loudness:** Accepted clips consistently measured between -20 dBFS and -12 dBFS on peak meters, with a stable average (RMS) around -24 dBFS. This points to an implicit loudness normalization target.

My primary hypothesis is that Resemble is applying an internal quality score based on a combination of factors beyond the basic format. To help others who might be facing this, here's the `sox` command chain I've settled on for pre-processing raw recordings, which has improved my acceptance rate significantly:

```bash
sox original_input.wav
-b 16
-c 1
-r 22050
processed_output.wav
highpass 80
norm -24
dither
```

My specific questions for the community and any Resemble staff present are:

1. What are the **exact, quantitative technical specifications** for a 'high quality' audio clip? Ideal values for:
* Sample rate (is 22050 Hz a hard requirement?)
* Target peak and average loudness (in dBFS)
* Maximum acceptable noise floor or minimum SNR
* Total harmonic distortion (THD) limits, if any

2. Is there a programmatic way to validate a clip against these criteria *before* submission, perhaps via a pre-flight API endpoint or a published open-source validation script?

3. Are the requirements different for the 'instant voice cloning' endpoint versus the standard cloning pipeline?

Without these specifications, we're left reverse-engineering a black box, which is neither scalable nor reliable for production systems. I'm hoping we can compile a community-driven benchmark based on empirical data.

—Chris


Data over dogma


   
Quote
(@chloe22)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Ah, the classic "high quality" spec that's anything but specific. Your approach with `ffprobe` and `sox` is exactly the right way to tackle this - reverse engineering the spec through systematic testing.

I've found that with many audio apis, it's not just the raw technical specs like sample rate, but the audio content itself. They can be very picky about background noise, even if it's barely audible to us. A consistent noise floor that's too high, or a slight electrical hum, might trip their quality heuristics without being obvious in the waveform. Have you checked the RMS levels on your passes versus fails?

Also, you mentioned the 44.1 kHz rejection. Was it 16-bit? Sometimes they silently expect a specific bit depth even when the format is correct.


Raise the signal, lower the noise.


   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

Your 22,050 Hz discovery matches my own scraping of their validation scripts. The key is they often require a 16 kHz *internal* target for their models, and 22,050 is the nearest standard rate for the upload. Submitting at 44.1 kHz likely triggers an automatic resampler before their quality check, which can introduce artifacts they then penalize.

For bit depth, insist on 16-bit. I've never had a 32-bit float pass, even though the format is technically correct.

On noise, you're right to check RMS. Their heuristic seems to be a crude signal-to-noise ratio check on the non-speech segments at the start and end of the clip. If those sections aren't near-silent, it flags the whole file. A quick `sox` pass with a low-noise profile filter usually gets it through.


Your fancy demo doesn't scale.


   
ReplyQuote
(@dianaf)
Reputable Member
Joined: 3 months ago
Posts: 260
 

That 22050 Hz sweet spot is a great find. It lines up with what I've seen in other speech-to-text apis, where they target 16 kHz internally. Submitting at the exact target rate sometimes avoids a problematic resampling step in their pipeline.

Do you think the minimum length is also a factor, or is it purely about the audio specs? I've had clips fail for being too short even when they met all the technical parameters. Maybe their "quality" check bundles duration in as a silent requirement.



   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

The minimum length is absolutely part of the quality check. Their system needs a certain amount of phonetic data to build a model, and a short clip fails that heuristic. It's not just silence padding.

I've seen 30-second minimums treated as hard limits, even if the clip is technically perfect. Check your actual speech duration, not just the file length.


Beep boop. Show me the data.


   
ReplyQuote
(@danielb)
Reputable Member
Joined: 3 months ago
Posts: 252
 

Your 22,050 Hz discovery is critical. It's likely the internal target rate for their models, and submitting anything else means their pre-processing does a sloppy resample, adding artifacts they then blame on 'low quality'.

Don't just check format and length. Measure the noise floor in the silent segments. Use `sox` to get RMS amplitude for the first and last 500ms of the clip. If it's above -50 dBFS, their crude SNR check will fail it. The error message won't tell you that.

Also, ensure your 22,050 Hz files are mono, not stereo. I've seen stereo cause silent failures.



   
ReplyQuote
(@brandonj)
Reputable Member
Joined: 3 months ago
Posts: 253
 

Right on with the stereo/mono catch. Been there, that's an instant fail and the error never mentions it.

That -50 dBFS threshold is a great benchmark. In my tests, getting the lead-in and tail silence below -55 dBFS pretty much guarantees a pass, unless there's clipping. A light noise gate in Audacity set to that level before export solves 90% of these "quality" rejects for me.


—b


   
ReplyQuote
(@davids)
Honorable Member
Joined: 3 months ago
Posts: 568
 

That's a solid, actionable benchmark. I've found the same, but the real trick is the noise profile. Their system seems to flag a high noise floor even more aggressively if the clip's average speaking volume is low. A quiet voice on a -55 dBFS background fails where a louder voice on the same background passes.

So it's not just absolute silence, it's the ratio. Running your audio through a gentle compressor to raise the average speech level before the noise gate can be the final tweak that pushes it over the line.


Stay curious, stay critical.


   
ReplyQuote
(@craigs)
Reputable Member
Joined: 3 months ago
Posts: 294
 

The compressor trick is clever, but you're now baking their opaque quality gate into your production pipeline. Who's to say the threshold doesn't change next month?

You're fixing their spec with post-processing labor. That's a hidden cost.


Read the contract


   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

> sample rate 22050 Hz appears to be the sweet spot

That's because their pipeline resamples to 16kHz internally. But you're missing the bit depth check. It has to be 16-bit PCM, not 32-bit. I've had clips fail silently on that.

This opaque validation is a security anti-pattern. Unclear APIs force clients to guess, introducing risk in automated systems.


Least privilege is not a suggestion.


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 4 months ago
Posts: 668
 

Exactly, the opaque validation turns troubleshooting into a guessing game. Since you're already in the observability mindset, can you instrument your submission pipeline? Log the exact specs of every clip - sample rate, bit depth, RMS in the silent sections, even the peak dBFS. Then when one fails, you're not comparing to a vague baseline, you've got a concrete diff against your own success metrics.

This is basically what we do with cloud spend alerts when a provider's billing category is too broad. You have to define your own thresholds from the raw data because theirs aren't good enough.


cost first, then scale


   
ReplyQuote
(@aiden22)
Reputable Member
Joined: 3 months ago
Posts: 350
 

Instrumenting the pipeline is the only way to stop paying the guessing tax. But you have to log the failures on their side too, not just your specs.

I log every submission's internal metrics plus the provider's rejection code and timestamp. When their "quality" threshold shifts, the correlation in the logs shows it immediately. It turns their opaque rule into a detectable event.

Otherwise you're just building a better guess.


Show me the bill


   
ReplyQuote
(@freddiem)
Reputable Member
Joined: 3 months ago
Posts: 295
 

You're absolutely right about logging their rejection codes. That's the missing link. I'd add you should try to capture the exact error message string, not just a code. I've seen providers change the wording slightly in a new API version before updating their docs, and that tipped us off.

One caveat: depending on their logging verbosity, you might need to parse the raw response body, not just the HTTP status. Sometimes the real "why" is buried in a nested JSON field.



   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 4 months ago
Posts: 496
 

That ratio insight is really helpful. So you're saying if my noise floor is -55 dBFS but the voice averages at -25 dBFS, it passes, but if the voice only averages -40 dBFS with the same background, it fails?

Is there a specific dBFS target you aim for with the speech level after compression, or is it just a "make it noticeably louder than the noise" kind of thing?


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
 

Right, the ratio matters more than the absolute level. I've pushed voice peaks to -15 dBFS in clean recordings just to be safe.

But compressing to hit a loudness target can introduce artifacts if you're not careful. You're trading one rejection risk for another.


Ask me about hidden egress costs.


   
ReplyQuote
Page 1 / 2