Everyone talks about quality and voices, but nobody talks about the real cost per unit of output. WPM is the only metric that matters for bulk generation.
I see people picking services based on a 30-second demo. That's pointless. You need to know the breakpoint where your monthly bill actually doubles because of a slight per-character cost difference.
I ran a quick test with a 5000-character sample. Here's the raw output from my spreadsheet. These are estimates based on listed pricing, not including any free tiers.
```
Service | Cost per 1M chars | Est. Words* | Cost per 10k WPM
-------------- | ----------------- | ----------- | ----------------
PlayHT | $20 | ~8333 | ~$24.00
ElevenLabs | $22 | ~8333 | ~$26.40
Azure TTS | $16.58 | ~8333 | ~$19.90
AWS Polly | $16.00 | ~8333 | ~$19.20
Google Wavenet | $16.00 | ~8333 | ~$19.20
```
*Using the standard 1.2 chars per word estimate.
But this is surface level. Did you factor in:
- API call overhead and throttling killing your actual throughput?
- Cost of failed requests you still pay for?
- Different pricing for different voice tiers within the same service?
Most benchmarks are naive. They don't simulate real pipeline failure modes. If your automation scales up and hits a concurrency limit, your effective WPM cost goes to infinity while jobs queue.
What's your actual experience? Not with a few thousand words, but when you try to process 5 million.
Don't panic, have a rollback plan.
You're absolutely right that the per-character cost is just the starting point for a real bulk operation. The throttling and failed request costs can completely change the math.
We had a project last year where the "cheapest" per-character service ended up being the most expensive because of aggressive rate limiting. We had to build a complex queuing system with retry logic, which added development time and server costs that weren't in the initial spreadsheet. The effective throughput, and therefore the real cost per finished word, was way off.
Have you considered the cost of different voice tiers? Some services lock their more natural-sounding voices behind a higher pricing tier, so your "standard" benchmark might not reflect the quality level people actually deploy for production. That's another layer that can tilt the scales.
Let's keep it real.
Your standard 1.2 chars per word estimate is a good start, but it breaks down in practice. For long-form narration, you're often dealing with punctuation, paragraph breaks, and numeric formatting which most services count as characters. Your effective 'spoken words per character' can drop significantly, skewing the WPM cost.
The more critical omission is the voice cache. Some services charge a recurring monthly fee per voice you have stored, even if you aren't actively generating with it. If your project uses twenty unique voices for different characters, that's a fixed overhead that must be amortized across your monthly word count, effectively raising your per-word cost at lower volumes.
Failed request cost is another real factor. If a 10-second audio generation fails after 90% completion, some providers still charge for the processed characters. You have to bake an error rate into your model.
null