Skip to content
Notifications
Clear all

Thoughts on the new 'concatenative' voice model they're hinting at?

3 Posts
3 Users
0 Reactions
6 Views
(@startup_selector)
Eminent Member
Joined: 4 months ago
Posts: 14
Topic starter   [#2367]

Hey everyone, been using WellSaid's free tier for a few months on my landing page videos. Saw a blog post about an upcoming "concatenative" voice model. They're saying it'll be more natural and flexible.

Can anyone explain what that actually means in practice? And more importantly, for someone on a budget, will this likely push the paid tiers up in price, or is it an upgrade to the existing voices? Really hoping it lands in the free tier or at least the starter plan. The current voices are good, but sometimes they still sound a bit... robotic in longer sentences.


Free tier is my favorite tier.


   
Quote
(@cost_analyst_ray)
Reputable Member
Joined: 5 months ago
Posts: 138
 

The term "concatenative" refers to an older, pre-deep-learning synthesis technique. It stitches together pre-recorded phoneme or word fragments from a human voice talent. It can sound very natural for short phrases, as you're hearing an actual human recording, but the flexibility is poor. Generating longer, novel sentences often leads to the robotic cadence you're noticing because the transitions between fragments are clumsy.

You're asking the right financial question. This isn't a next-generation model; it's a different, often more expensive-to-produce approach requiring massive recorded libraries. If they're marketing it as an upcoming feature, it's almost certainly a new premium product tier. The R&D and licensing costs for the voice talent will be passed on. I'd expect it as a paid add-on or part of a new, higher-cost plan. The existing neural model tiers will likely stay the same price, but they won't give this away for free. Budget for a 20-30% increase if you need the "natural" sound.


CostCutter


   
ReplyQuote
(@startup_ops_lead_alex)
Eminent Member
Joined: 1 month ago
Posts: 16
 

So you're saying the cost is high because they need a huge library of pre-recorded bits from a voice actor. That makes sense.

I'm curious, though, if the high production cost is the main reason they'd put it on a higher tier. Couldn't they also use it as a loss-leader to get people like me, on the free tier, to finally upgrade? If it's noticeably better for short clips, that's exactly what I use for intros and CTAs.



   
ReplyQuote