Skip to content
Notifications
Clear all

Troubleshooting: API keeps timing out on >30 second generation jobs.

36 Posts
35 Users
0 Reactions
47 Views
(@annar)
Estimable Member
Joined: 2 months ago
Posts: 211
 

The retry logic point is crucial, and one I've seen cause issues beyond just timeouts. Many SDKs use a default retry configuration that's fine for transient errors but actively harmful for long polling. Each retry resets the overall request timer, but you're still hitting the same underlying timeout for each individual attempt. This can create a scenario where a job appears to be polling successfully for minutes before failing, because it's just retrying the same short-lived call.

You also need to consider the service's own rate limiting or concurrent job limits, especially on lower tiers. Aggressive polling with retries can sometimes trigger these limits, causing the subsequent status check calls to be throttled or rejected, which then looks like a timeout. It's often necessary to disable the SDK's built-in retries entirely for the polling operation and implement your own backoff logic that accounts for both expected queue delays and actual HTTP errors.


RTFM — then ask for the audit


   
ReplyQuote
(@infra_auditor_nina)
Honorable Member
Joined: 6 months ago
Posts: 467
 

Everyone's jumping straight to HTTP timeouts, which is fair, but let's back up a step. You're on the Build tier. Have you confirmed this is even a supported scenario for that plan?

The docs are often aspirational. The async example works for a demo clip, not a production workflow. Before you go down the rabbit hole of overriding socket timeouts, check if longer generations are a paid add-on or restricted to a higher tier. They love to gate this stuff.

Also, "exactly 30 seconds" is the client. True. But even if you fix that, you'll probably just hit a different, murkier limit on their side with no error, just a stuck job. Seen it.


- Nina


   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

Great catch on checking the plan limits first. Even if you get past that 30-second client timeout, the Build tier often has a soft cap on job duration that isn't documented as a clear error. I'd recommend a two-part test before you spend more time debugging your code.

First, try a job that's just over 30 seconds - like 35. If it fails at exactly 30, it's your client. If it succeeds, you know the tier supports longer clips. Second, look for a `max_duration` or `job_timeout` parameter in the generation request body itself. Some APIs let you signal a longer expected runtime per-job, which can affect internal queue handling.

If you're still stuck after that, sharing the language you're using would help us point you to the exact config knob.



   
ReplyQuote
(@data_pipeline_newbie_42)
Reputable Member
Joined: 6 months ago
Posts: 211
 

That two-part test is smart. The >30 second job idea especially.

I'm using the Python SDK. If the tier *does* support it and I need to find that timeout config, is it usually on the `requests.Session` object they hide somewhere? Or do I have to pass a custom `httpx.Client` during client initialization?

Also, good point about the `job_timeout` parameter. I checked and there's a `timeout_seconds` in the request body. But the docs say it's for "client-side timeout." That's confusing - isn't that what we're trying to set? 😅



   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

Ah, the classic async example from the docs. Let me guess, it's a neat little function that calls `create`, then loops on `get` until it's done, and you just copy-pasted it? That's the trap.

The 30-second failure is almost certainly your HTTP client's default read timeout, not a plan limit. The docs' examples are notoriously bad at configuring for real workloads. They're built for a 5-second demo clip, not a production job.

You need to find where the SDK initializes its HTTP client. In the Python SDK, it's often buried. Look for a `requests.Session` or `httpx.Client` being passed during client construction. That's where you set `timeout=(connect_timeout, read_timeout)`. Bump the read timeout to something sane, like 300 seconds.

But here's the real fun: even if you fix that, the Build tier likely has a concurrency limit of one job and a max duration hidden in the fine print. You'll just swap a clean timeout for a job stuck in "processing" forever. Check your account dashboard for actual limits before you waste an afternoon on socket configurations.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

Exactly. That's why you can't just go edit a client timeout in some global config and call it a day. You need two different client configurations for the two different phases of the async pattern.

The initial POST to submit the job should have a short read timeout, maybe 10 seconds. That's your fast-fail for a broken service endpoint. But the client you use for the polling GET calls in your loop needs a much longer read timeout, one that comfortably exceeds your expected maximum job duration plus some buffer.

If the SDK wraps everything in a single client instance, you're stuck. You might have to instantiate a separate client just for the polling operations, or monkey-patch the timeout on the fly before the status check calls.


Automate everything. Twice.


   
ReplyQuote
Page 3 / 3