Skip to content
Notifications
Clear all

My results after stress-testing the API with 10k concurrent requests.

2 Posts
2 Users
0 Reactions
32 Views
(@consultant_carl_42)
Reputable Member
Joined: 4 months ago
Posts: 381
Topic starter   [#2983]

Alright, let’s get this out of the way before the usual chorus of “just use Murf, it’s scalable and modern” starts up. I was asked by a client to evaluate Murf’s enterprise readiness for a campaign that would involve burst traffic—think large-scale personalized outreach with voice synthesis. Their sales team heard somewhere that Murf could handle “any volume,” which is the kind of red flag that sends me straight for the API documentation and a stress-testing script.

I’m not here to talk about voice quality or the UI. I’m here to talk about what happens when you push the infrastructure. We simulated 10,000 concurrent requests to the text-to-speech generation endpoint, aiming to see where the cracks appeared. This wasn’t a casual curl command; this was a distributed load test from three different regions over a 15-minute window.

Here’s what you won’t find in the pricing sheet or the sales demo:

* **The 429 (Too Many Requests) wall is abrupt and poorly documented.** The published rate limits are… optimistic. In practice, we hit throttling at around 800-1000 requests per minute per API key, far earlier than we expected for an “Enterprise” tier. The error message is generic, and the `Retry-After` header was inconsistent—sometimes 30 seconds, sometimes 120. For a bursty marketing operation, this is a planning nightmare.
* **Concurrency is not the same as queue depth.** Murf’s system seems to accept requests happily, but the actual processing queue is a black box. We observed significant latency spikes (from 1.5s average to over 45s) once we crossed a certain threshold, even before explicit rate limit errors. Your client-facing application might think it’s fine, while jobs are silently backing up.
* **Cost implications of retries and failures are a hidden sinkhole.** When you get a 429 or a 5xx error (yes, we saw a few 502s under load), you have to decide: do you retry? If you’re building a resilient system, you will. That means for a burst of 10k, you might be making 15k-20k API calls due to retries, and you’re paying for every attempt, successful or not. Their billing is per request, not per successful synthesis.
* **The “scalability” they sell is vertical, not horizontal.** Adding more users or projects doesn’t inherently increase your API throughput ceiling. You need to negotiate specific rate limit increases and potentially pay for dedicated infrastructure, which is a sales conversation, not a checkbox in your account settings.

The takeaway? If you’re planning anything with a predictable, high-volume burst—launch-day campaigns, large batch personalization—you cannot rely on Murf’s standard API promises. You must:

1. Engage their enterprise sales team *before* signing.
2. Get explicit, contractual guarantees on rate limits and SLA for availability/latency under load.
3. Build significant client-side buffering, exponential backoff, and a fallback mechanism into your integration.

Otherwise, you’re building your castle on a foundation of “maybe.” I’ve seen multimillion-dollar campaigns stall for less.

-- Carl


Test the migration.


   
Quote
(@emilyr)
Reputable Member
Joined: 3 months ago
Posts: 295
 

That abrupt throttle behavior you describe is telling. I've observed similar patterns where the published rate limits assume an ideal, evenly distributed load, but real-world burst scenarios expose the underlying queueing or resource allocation logic. Did you capture the response headers on those 429s? The presence and values of `Retry-After` or `X-RateLimit-*` headers can reveal if it's a true, configurable API gateway limit or a symptom of backend saturation.

Your point about the error message being generic is a significant operational concern. For proper incident response and automation, you need granular error codes or at least a distinct `X-Cause` header to differentiate between, say, a global service quota exhaustion and a per-key limit. Without that, your client's SRE team would be flying blind during an actual surge.

What was your latency degradation profile leading up to the throttle? A gradual increase in P99 latency before a hard 429 is often more manageable than a sudden cliff, as it gives auto-scaling or alerting systems a chance to react.



   
ReplyQuote