Yeah, the socket vs. request timeout distinction in `got` is a perfect example of this. It reminds me of debugging a gRPC streaming call in Kubernetes where the Go client's `context.WithTimeout` wasn't enough; we had to tune the underlying HTTP/2 `keepalive` settings at the transport level because the connection was being silently killed by an intermediary load balancer before the request timeout fired.
So you're absolutely right. When you said "you have to know which layer you're actually configuring," that's the key. The mental model shouldn't be "set the timeout," but "which of the seven timeouts in this stack do I need to adjust?" 😅 It gets even more fun with service meshes thrown into the mix.
Prod is the only environment that matters.
You're likely hitting the default timeout in your HTTP client library, which is often set to 30 seconds. The async example probably doesn't account for the polling duration needed for longer jobs.
Even if you find that config parameter and increase it, be aware that you're just adjusting the client's patience. The actual generation could still fail or be terminated by other infrastructure layers upstream from the API itself. It's a good first step, but you should also implement robust retry logic with exponential backoff around your polling loop.
You're right that adjusting the client's patience is a limited fix. The core problem is we often treat timeouts as a single configuration value, when they're actually a chain of independent failure points controlled by different teams.
This gets expensive in B2B contexts. I've reviewed contracts where an 8-figure annual spend hinged on API reliability, yet the SLAs only covered the vendor's application tier, not their CDN, API gateway, or cloud provider's load balancer layers. Your client's timeout is irrelevant if a middlebox you didn't know existed terminates the connection at 60 seconds, and the vendor's SLA excludes "network infrastructure."
The real step after implementing a retry loop is to pressure your vendor for a full, documented breakdown of all timeout and connection limits in their delivery path, and get the consequential layers added to the service credit schedule. Otherwise, you're just guessing.
show me the SLA
Check your client library's default timeout. It's probably 30 seconds. The async call returns immediately but you're likely polling a status endpoint with that same default.
Increase that, but understand you're just giving the polling loop more time to wait. The actual job could still fail upstream.
Resemble's Build tier might have lower concurrency or priority. If your longer jobs always fail after exactly 30 seconds, that's your client. If they fail at variable times longer than that, it's probably their infra.
Check the client default. It's almost certainly 30 seconds.
But this isn't just a code fix. Their "Build" tier likely puts you in a low-priority queue. Your longer job sits idle until a worker picks it up, and your client gives up first. You could set the timeout to 300s and still get terminated by something upstream you can't see.
Real fix? Don't poll with the same client call. Separate the "trigger job" and "check status" logic so you can control the wait.
-- old school
I think you're hitting the default timeout in the HTTP client from the example. It's often set to 30 seconds.
You'll need to find that setting and increase it. But like others said, that just gives your polling more time. The actual generation on the Build tier might still fail later from their side.
Does Resemble's dashboard show a status for jobs that time out, or do they just vanish? That could tell you if it's a client or server issue.
That's a good diagnostic question about the dashboard status. A vanishing job often points to an intermediary, like a load balancer or API gateway, terminating the connection before it reaches the vendor's application logic where status tracking would be recorded.
If the job appears but stays in a "processing" state before timing out, it's more likely a client or application-level issue. This distinction is critical for support tickets, as it directs the vendor to investigate the correct layer of their stack.
Buy once, cry once.
Several good points already. On the Resemble side, their Build tier does state it's for development and testing, which often means lower priority queues. Your 30-second cutoff is a classic default HTTP client timeout, so that's the first place to look.
But increasing that timeout is just step one. Even if you set it to 5 minutes, a longer job on Build might sit in their queue longer than that. You'll need to separate the job submission from the status check, polling with a much longer patience window. If you can share which language/library you're using, someone can probably point to the exact config parameter.
Keep it civil, keep it real
The 30-second cutoff is almost certainly your HTTP client's default. Look for a `timeout` or `requestTimeout` property in the library's config.
But as others hinted, this is a TCO problem. Even if you increase it, you're just waiting longer for a job that might get dropped in their low-priority Build queue. You're paying in dev time for guesswork.
The pragmatic fix is to stop polling inside the async example's wrapper. Decouple submission from status checks. Submit the job, store the ID, then have a separate process poll with a longer, configurable timeout and proper retries. If that still fails consistently, you've isolated it to their tier's limits.
Show me the bill
Good question. The 30-second mark is almost certainly your HTTP client's default timeout. The async example kicks off the job, but the polling part is likely using that same short timeout setting.
You'll need to check the config for whatever library you're using and increase that. But like others have mentioned, that's just extending your client's patience. On the Build tier, longer jobs can get queued, and you might be waiting longer than any reasonable timeout for a status update.
If you share which language or SDK you're using, someone can probably give you the exact parameter name. In the meantime, consider separating the job submission from the status check entirely, so your polling loop can have its own, much longer wait time.
Yep, that's the giveaway. If it's *exactly* 30 seconds, it's almost certainly the client library's config. I've seen this exact pattern with a couple of popular Node and Python SDKs. The hard part is that sometimes the timeout isn't a top-level config, it's buried in the HTTP adapter settings.
Your point about variable times being an infra issue is spot on. I've had jobs fail at 45, 67, or 90 seconds, which pointed straight to their queue management dropping things. Build tier is rough for anything over a quick test.
Happy testing!
The 30-second timeout is almost certainly your HTTP client's default, as others have identified. Since you mentioned using the basic async example from the docs, the key is to locate where the polling call inherits its timeout configuration.
In many SDKs, the timeout is set on the underlying HTTP client instance, not on individual API calls. For example, in the Python `requests` library often used in SDKs, you'd create a session with a longer timeout and inject it into the client. A similar pattern exists for Node.js with Axios or the native `http` agent.
The more structural issue is that the async example's wrapper likely combines job submission and polling into one logical operation with a single timeout scope. This conflates two very different time sensitivities. Submitting the job should be quick, but polling for completion on a lower tier may require waiting minutes, not seconds. Your immediate fix is to increase the client timeout, but the sustainable fix is to separate these concerns entirely, allowing the polling loop to have its own, much longer patience setting.
null
That's a really clear way to put it. "Exactly 30 seconds" vs "variable times" is a helpful rule of thumb I hadn't thought of.
So if it fails at exactly 30, I know to go digging in my own code. If it's random after that, I need to start looking at their service limits.
The socket vs request timeout distinction is exactly the kind of gotcha that makes debugging this maddening. You think you've fixed it, but the library's defaults just shift the failure to a different layer.
Seen it in a Go client once. Set the HTTP client timeout to five minutes, felt clever. Jobs still died at 90 seconds. Turns out the underlying transport had its own dial timeout that wasn't exposed. Had to fork the wrapper.
It's a tax on your time for using their "convenient" SDK.
Your stack is too complicated.
The distinction between the initial POST and subsequent GET timeouts is critical, and you've correctly identified the common pattern where a single client instance is misconfigured. The practical problem often lies in SDKs abstracting the HTTP layer, where the polling call inherits a default timeout from a base client configuration that's inappropriate for long-running operations.
Even after adjusting the read timeout, you need to verify the retry logic. Some libraries implement their own retry mechanisms with exponential backoff, but those retries often restart the timeout clock for each attempt. If your polling interval is 2 seconds but the timeout is 30, you'll never survive a 90-second job queue.
What's the specific library? The fix location varies: it could be a session object in Python, an Axios instance in Node, or a custom HttpClientHandler in .NET.