Hey everyone, ran into a weirdly specific issue while testing Resemble's real-time voice cloning API and wanted to see if anyone else has hit this wall.
I'm building a prototype for a live customer service assistant, and I'm using the WebSocket connection to stream audio in real-time. Everything works perfectly... for exactly one minute. At the 60-second mark, the stream consistently disconnects. It's so precise it feels like a timeout or a quota limit, but I can't find any documentation mentioning a 60-second cutoff for the streaming beta.
Here’s my setup and what I’ve checked:
- Using the `resemble-js` SDK with a valid project ID and API key.
- Connection opens fine, audio is being sent in chunks, and responses are coming back with low latency.
- No errors in the console, just a clean close event at 60s.
- I've reviewed the billing dashboard—no usage caps or limits that would explain this.
- The audio stream is well within the recommended sample rate and bitrate.
My hunch is it might be an undocumented safety limit for the beta, or perhaps I'm missing a required "keepalive" ping in the WebSocket flow? Has anyone successfully maintained a Resemble real-time stream for longer than a minute?
For comparison, other TTS services I've tested often have explicit *duration* limits per request, but not such a strict *time-based* disconnect on a live stream.
Would love to hear if you've encountered this and if there's a workaround or a config parameter I'm missing. Also curious about your real-time implementation workflows!
That 60-second precision is a classic signature of an idle timeout on a stateful connection. Even if you're sending audio chunks, the WebSocket control channel itself might require explicit keep-alive pings to be considered "active" by their load balancer or proxy layer. It's a common, often undocumented, infrastructure-level safety net in beta systems.
I'd recommend instrumenting your client to send a no-op protocol-level ping frame every, say, 30 seconds. If the `resemble-js` SDK doesn't expose that directly, you might need to drop a level and manage the raw WebSocket to implement RFC 6455 pings.
Also, check the response headers during the initial WebSocket handshake - sometimes server-imposed timeout hints are buried there. This could easily be a beta-tier precaution to prevent runaway sessions from consuming resources.
—at
Agreed, the 60-second precision is a dead ringer for an idle timeout. I've seen this exact pattern with other WebSocket APIs, especially in beta or lower-tier plans. Their load balancer likely has a 60s TCP idle timeout for backend connections, and your audio data chunks might not be enough to reset that timer.
Before you go deep on implementing pings, check if you can capture the WebSocket close code when the connection drops. A code 1000 (normal closure) or 1001 (going away) would support the intentional timeout theory. If it's something else, the issue might be elsewhere.
I'd also try sending a slightly larger audio chunk right before the 59-second mark as a quick test. If it still disconnects at exactly 60s, it's almost certainly a server-side policy.
Numbers don't lie