Alright, gather ‘round the digital campfire for another installment of “What Actually Works?” This time, it’s Resemble AI’s voice cloning in the hot seat. We’ve been integrating synthetic voices into a customer-facing chatbot for about 18 months now, having previously burned through a couple of other providers (whose names rhyme with “Bleven Laps” and “Moorable”). The promise of a seamless, real-time conversation is the holy grail, right? So we had to test the two main delivery methods Resemble offers: the so-called ‘Real-time’ API and the ‘Async’ (batch) API. The marketing gloss suggests the former is for interactive use and the latter for bulk generation, but the reality, as always, is in the latency and cost trenches.
Let’s cut to the data. We ran 500 identical short prompts (typical chatbot responses, averaging ~120 characters) through both endpoints, using the same voice clone. We measured from the moment our system sent the POST request to the moment we received the final, usable audio file URL. Network variance was minimized. Here’s the grim—or perhaps just realistic—picture:
**Real-time API Results:**
* **Average Latency:** 1.8 seconds. Let’s be clear: in a live chat interaction, waiting nearly two seconds for the *audio* to start playing *after* the text is ready is a lifetime. The user experience isn’t “real-time”; it’s “noticeable pause.”
* **Cost:** Priced per second of generated audio. Our tests averaged about 3.5 seconds of audio per request. The per-request cost was predictable but adds up fast in a high-volume chat.
* **The Kicker:** The ‘real-time’ endpoint still doesn’t stream the audio chunk-by-chunk as it generates. You’re waiting for the entire file to be rendered on their servers before you get anything back. This is just a faster batch process, not true streaming synthesis.
**Async API Results:**
* **Average Latency:** 4.7 seconds. Obviously slower, but here’s the thing: for our use case, it was often *functionally equivalent*. Why? Because our chatbot workflow already involves some processing time to formulate the response text. By the time we had the text ready, we could fire off the Async request and the total added wait was often marginal.
* **Cost:** Significantly cheaper. The per-second rate is lower, and for some volume plans, the cost advantage is substantial.
* **Hidden Benefit:** Reliability. We observed fewer timeouts and transient errors with the Async endpoint during peak loads. The ‘real-time’ endpoint, under stress, would occasionally just drop a request, forcing a retry and making the latency even worse.
So, what did we learn? The ‘real-time’ label feels like a misnomer, or at least a generous interpretation. Unless you have a use case where every millisecond counts *and* you can stomach the premium, the Async API might be the sneaky-smart choice. We’ve actually re-architected our chatbot to use the Async API for all non-urgent responses and keep a cache of commonly generated phrases. The cost savings are notable, and the perceived latency for the end-user is virtually unchanged from the ‘real-time’ option. It feels like paying for premium gasoline when your engine is tuned for regular. The real question for Resemble is: when will we see a *truly* real-time, streaming API that can start delivering audio chunks before the entire sentence is synthesized? That’s the only way this makes sense for live conversation. Until then, the Async API is the contrarian’s choice.
I'm the sole bookkeeper at a 12-person e-commerce shop. We use Resemble's Async API in prod to generate post-sale customer service updates voiced by our founder's clone.
**Key differences from our testing:**
**Cost per clip:** Async is about 30% cheaper per second of audio generated. Real-time's premium is for the delivery model, not higher quality.
**True latency:** Async takes 5-7 seconds for those short clips to be available for download. Real-time was 1.8-2.5 seconds, but that's still not "conversational" for us.
**Architecture fit:** Real-time requires keeping a WebSocket open, which was a pain point with our serverless chatbot backend. Async just uses a simple webhook callback.
**Error handling:** Failed Real-time generation just drops the socket. Async gives us a retry hook and a status API, which is more reliable for our logs.
My pick is the Async API, unless your chatbot absolutely requires sub-2-second audio delivery and you can handle the socket management. Tell us your current peak concurrent user load and if you're using serverless, and I can get more specific.
I've been down a similar road with a different provider's "real-time" endpoint. My experience lines up with yours - 1.8 seconds felt like an eternity in a chat flow where users expect sub-second replies. The real gotcha for us was that the streaming start wasn't actually instant either; the first byte took almost as long as the full async response.
One thing you didn't mention: did you measure the time to first audio chunk vs full audio available? For our use case (turn-by-turn chat), we needed to start playing the response while it was still generating. Neither API really supports that smoothly unless you're willing to buffer aggressively.
Also, I'm curious - did you see any difference in voice quality between the two modes? I've heard rumors that async uses a slightly heavier model, but my own tests were inconclusive.
You're spot on about the first-byte latency being the critical metric for turn-by-turn chat. In our tests, the real-time API's time to first chunk was indeed nearly identical to the full clip generation time, around 1.8 seconds. The streaming protocol feels more about efficient transmission after the fact, not progressive generation.
On voice quality, we did run spectrogram analysis and couldn't find a statistically significant difference between the outputs of the two APIs for the same voice model. The rumor about async using a "heavier" model might stem from providers using a more batch-optimized, non-streamable architecture, but the perceptual quality seems identical.
The buffering strategy you mention is the only workable path. We ended up implementing a 500ms pre-buffer on the client before playback starts, which masks a portion of the initial latency, but it's a trade-off against perceived responsiveness.
CPU cycles matter
The spectrogram analysis is a solid approach. We performed a similar comparative analysis last quarter using PESQ and POLQA scores on a corpus of 1,000 synthetic utterances, and the results aligned with yours - no meaningful perceptual difference. This strongly suggests the underlying acoustic model is identical.
However, the >time to first chunk was nearly identical to the full clip generation time< is the critical failure mode for any "real-time" claim in an interactive setting. It indicates the system is performing full non-streaming inference before packetization begins. The WebSocket is merely a transport layer optimization.
Our workaround was more aggressive: we implemented a speculative generation cache. When the user's intent confidence exceeds a threshold, we trigger async generation for the three most probable next system responses in the background. If we guess correctly, the audio is ready instantly from cache; if not, we fall back to the standard API with the usual penalty. This shifted our P99 latency for correct predictions from 1800ms to under 50ms. The cost of wrong predictions is, of course, higher.
The speculative cache is a clever hack. We tried something similar but ran into a different cost issue: audio storage. When you're pre-generating multiple potential responses for thousands of concurrent sessions, the S3 bills added up fast, especially for longer clips. Did you factor that into your total cost of ownership?
Also, on your point about the WebSocket just being transport, that matches what we saw. Is there any actual streaming model on the market, or are they all doing full inference first?
The storage cost is a brutal hidden tax on speculative caching. Our solution was to pre-generate only the first 500ms of audio for high-confidence intents - just enough to mask the initial latency - and let the rest stream live. It reduced our storage footprint by about 80% compared to full clips.
>Is there any actual streaming model on the market?
Not in the true sense of progressive token generation, no. Every vendor we've tested, including the big ones, does full server-side inference before sending a byte. The "real-time" label is, charitably, marketing for "low-latency batch with a fancy pipe". The only genuine streaming we've built used a locally-hosted, heavily-optimized model, and that's a whole other horror story of GPU debt.
APIs are not magic.
The partial pre-generation is a smart mitigation, but it introduces a new failure mode: what happens when the live stream for the rest of the clip fails or is delayed? You're now managing a hybrid system.
On the GPU debt horror story, that's the real cost. Running even an "optimized" local model for true streaming requires a reserved, high-memory instance. The cloud API markup starts to look reasonable once you factor in engineering time and the constant risk of OOM kills during peak concurrency.
Your fancy demo doesn't scale.