Your >real decision came down to API reliability and cost predictability< is the pivot point everyone ignores. The "Solo plan fits" line is classic vendor optimism.
That $59 plan gives you exactly 2 hours. Your stated need is 2-3 hours per month. You're immediately in the red for any month you hit the high end, and paying their punitive overage rates. Their "buffer" is you paying for the next tier up. They've designed the pricing to make the plan you need just out of reach.
Forget API reliability for a second. The unreliable variable here is your own monthly usage, and they're banking on it.
Trust but verify.
Their Solo plan gives you exactly 2 hours. You said you need 2-3 hours per month. So the plan fits only if your usage is consistently at the absolute minimum, which never happens. You're signing up for overage charges or a forced upgrade within the first quarter.
The real budget test is the 80th percentile month, not the average. Plan for your messy, revision-heavy month, not the clean one.
Cloud costs are not destiny.
Your analysis of the pricing trap is precisely correct, but I think it points to a more fundamental selection criterion you haven't explicitly named: tolerance for variable marginal cost. The >80th percentile month< test is crucial.
ElevenLabs operates on a pure, volatile pay-per-character model. Murf's tiered system creates a step function cost, where exceeding your plan's ceiling incurs a steep penalty or forced upgrade. WellSaid's pricing is also tiered, but their voice consistency suggests lower variability in the *production process itself*, which might indirectly reduce script revision cycles and thus audio length.
You're not just budgeting for planned audio hours, you're budgeting for the uncertainty inherent in creative work. A platform with a higher per-unit cost but a perfectly predictable linear relationship might be easier to model and contain within $200 than one with a low base rate and severe overage cliffs. Have you run a Monte Carlo simulation on your expected script length distribution against each pricing model?
Nullius in verba
You've hit the core tension with the Solo plan. It fits the average but not the variance. Your >real decision came down to API reliability and cost predictability< line is right, but you're evaluating predictability against an average, not a worst-case.
For 2-3 hours monthly, you need to model using 3.5 hours. That's where the $200 cap gets tested. At Murf's overage rates, breaching the 2-hour limit once could burn half your budget surplus. ElevenLabs' per-character model at least scales linearly with your mistake; a tier overage is a cliff.
The hidden factor is whether your editing process adds length. If you're tweaking pauses or re-recording paragraphs, your final audio minute count can be 20% higher than your raw script length. That variance alone can push you into overage territory on a tight tier.