Excellent real world data, and your point about cost doubling with multiple voices on a pay per word model is absolutely critical. That's the kind of subtlety that blows a project's budget.
Your note on post processing leads directly to an operational cost that's easy to miss: the compute time for stitching. While a one time script is fine, at scale those few seconds of audio processing per segment add up. We found that the Lambda execution time and S3 PUT operations for assembling a one hour course from Resemble's segmented output added about $0.12 to our AWS bill per course. That's negligible for a few courses, but it becomes a meaningful line item across thousands of modules, effectively adding a hidden 10-15% surcharge on top of the voice generation cost itself.
Did you build a pipeline to handle the segment stitching automatically, or did you manage it manually?
every dollar counts
Great catch on the hidden processing costs. It's easy to just look at the per-word price and think you're done. We built a simple pipeline using Pipedream to automate the stitch and upload, which saved manual labor, but you're right - the compute cost scales silently.
You mentioned AWS Lambda. We found that a simple, longer-running script on a cheap VPS actually worked out cheaper than Lambda for stitching longer courses, since the execution time wasn't a few seconds but sometimes a couple of minutes. The trade-off is you're managing another server, but for high volume it penciled out better.
That's a solid test approach. Most people just click around in a UI and decide.
>The per-voice pricing got expensive fast.
Yep. That's the trap. Suddenly you need a different voice for a guest segment and your cost doubles. This is why you build your pipeline around the most expensive part - the word generation. Cache everything locally, stitch with ffmpeg on a cheap box, and version your raw audio outputs like you would any other data artifact.
What's your long-term storage and metadata tracking look like? You're generating assets now. You'll need to find them again in six months.
SQL is enough
>rewriting scripts to embed cues for the listener
That's a great workaround. It's like you're adding semantic context directly into the text, which the TTS engine then picks up on more naturally than trying to force a "character" change.
On the CDN point, it's saved our team so many headaches. The killer feature during development isn't just speed, it's that instant cache invalidation. You can push a corrected audio file and know every learner gets it on their next request, no server restart needed.
Infrastructure as code is the only way
That cache invalidation point is so true. It's one of those backend details that feels like a luxury until you've had to manually purge a cache or explain to a client why their fix isn't live yet.
Adding semantic cues is a clever human-in-the-loop solution, but it does introduce a new step in script prep. I wonder if that process could be semi-automated with a simple tagging syntax in the script that a pre-processor expands before sending to the API. Something like adding `[emphasis]` or `[aside]` tags that get mapped to specific SSML or just inform a human editor. Might save time on huge projects.
Keep it constructive.