Hi everyone. I've been seeing more folks in the community talk about scaling up their Synthesia projects, moving from one-off videos to larger batches. The promise of the API for automating this is clear, but I'm curious about the practical, real-world experience.
Specifically, for those who have moved beyond simple test calls: what's the actual latency like when you're generating, say, 50 or 100 videos in a queue? I'm not just talking about the time-to-first-byte, but the total end-to-end clock time for a sizable batch. Does the system throttle or queue requests in a noticeable way? Are there any patterns you've noticed (e.g., faster at certain times, slower with certain avatars or languages)?
I'm also interested in how you structured your calls for reliability. Did you implement a specific polling interval, error handling, or parallelization strategy that worked well? Understanding these operational details would help a lot of us plan our workflows better.
Looking forward to hearing about your hands-on experiences and any gotchas you might have encountered.
Keep it civil, keep it real.
Hey user622, that's a great set of questions. The latency for large batches can be surprisingly variable in my experience. For a queue of 100, total clock time really depends on current load, and you'll often see the system start to queue requests internally after the first 10-15 go through. I've noticed it tends to handle them in waves.
On structure, we implemented a simple exponential backoff for polling the status endpoint. A fixed interval will burn you if one video gets stuck. The bigger gotcha is that some avatar/language combos do take noticeably longer, so if your batch isn't uniform, that throws off your total time estimate.
Good, practical questions. The latency you're asking about is definitely a real factor in planning. I'd add that from what we've seen, the queue depth and total batch time aren't linear. Submitting 50 videos doesn't take 5x longer than 10, it can be 8-10x due to that internal queueing user1172 mentioned.
For structuring calls, a backoff strategy for polling is key. But I'd also stress the importance of monitoring *individual* video statuses separately within your batch. If one fails or gets stuck, you need to be able to isolate it without halting your entire pipeline. Some folks build in a checkpoint system to log each video's progress independently.
Have you settled on a specific language or set of avatars for your project? That can really change the ballpark numbers.
Keep it constructive.
Great questions. Following along because we're planning a bulk run soon too, and timing is crucial for our internal deadlines.
A point I'm still trying to nail down: for the *total clock time*, does anyone have a sense of how much the script or polling overhead adds versus pure render time? In your 100 video batches, was most of the delay just the API queue, or did your own backoff logic stretch it out significantly?
Also, have you noticed any difference in queuing behavior between submitting all 100 requests at once versus a staggered approach, say, 10 every minute?
The latency is the easy part to measure. The real question is if that latency is consistent enough to depend on for any automated workflow with a deadline. My experience says it's not.
You can implement all the backoff and parallelization you want, but the queue behavior is opaque. I've seen a 50-video batch finish in an hour one day and take four the next with no change in our script or the content. Calling it a 'queue' suggests a predictable order, which is generous.
Your point about specific avatars and languages is where the marketing slides meet the messy reality. Some of those 'premium' avatars seem to run on a single, very tired server somewhere. If your batch isn't perfectly uniform, forget about a reliable total time estimate.
Trust but verify