Hi everyone. I'm pretty new to this whole B2B software evaluation thing, so please bear with me. I just wanted to share my small team's experience switching our text-to-speech provider, because it's been a bit of a mixed bag.
We were using IBM Watson Text to Speech for about a year. The voices were good, but honestly, the whole setup felt really complex. The API documentation was huge, and we always felt we were only using a tiny fraction of what was possible. It was powerful, but maybe *too* powerful for what we needed—just generating voiceovers for product explainer videos and training modules.
Last month, we switched to WellSaid Labs. The difference is night and day in terms of ease of use. Their studio is so straightforward. We can type in a script, pick a voice, and have a high-quality MP3 in minutes. No messing with SSML tags or worrying about complex API calls. The team loves it because anyone can jump in and create something.
But here's my question, and maybe some of you with more experience can chime in. While WellSaid is easier, I sometimes feel like we lost something. With Watson, we could fine-tune pronunciation, control pitch and speaking rate very precisely, and even train custom voices (though that was way out of our budget). It felt like we were tapping into decades of R&D. With WellSaid, we get amazing out-of-the-box quality, but less fine-grained control.
Is this just the trade-off? Simpler tool equals less advanced controls? For those who've used both, do you ever go back to a more "heavy" solution like Watson for specific projects, or is the convenience of WellSaid worth sticking with for everything? Really curious to learn from your workflows.