Skip to content
Notifications
Clear all

Has anyone used WellSaid for generating audio for visually impaired users on a website?

3 Posts
3 Users
0 Reactions
23 Views
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
Topic starter   [#22591]

Greetings, community members. This is a query that sits at an interesting intersection of accessibility, user experience, and synthetic voice technology, and I believe it warrants a detailed discussion. While we often examine text-to-speech (TTS) platforms like WellSaid Labs through the lens of content creation, marketing, or e-learning, their application in making web content accessible to visually impaired users presents a distinct set of criteria and potential challenges.

The core question is whether a high-quality, proprietary TTS service like WellSaid is a suitable and practical solution for generating real-time or pre-rendered audio versions of website content. Key considerations here would include the technical implementation, cost structure relative to the scale of content, and most critically, how the output aligns with established web accessibility standards (WCAG). For instance, does the natural cadence and intonation of a WellSaid voice agent enhance comprehension over a standard system voice, and does that justify the operational complexity? Furthermore, one must consider control elements: can a user pause, rewind, or adjust the speech rate, which are common expectations in screen reader software?

I am particularly interested in hearing from members who have implemented or piloted such a system. Practical insights would be invaluable. For example:
* Was the audio generated dynamically via API for each page load, or was it pre-rendered for static content?
* How did you handle frequently updated content, and what was the associated latency or cost implication?
* Did you conduct any user testing with visually impaired individuals to gauge the effectiveness and usability compared to traditional screen readers like JAWS, NVDA, or VoiceOver?
* Were there any specific technical hurdles in embedding the audio players in an accessible manner (e.g., proper ARIA labels, keyboard navigation)?

The philosophical and practical balance here is fascinating—between leveraging cutting-edge, human-like TTS for a better experience and ensuring the solution remains robust, affordable, and, above all, functionally accessible. I look forward to a constructive exchange of experiences and technical perspectives on this nuanced topic.

— EthanP


Let's keep it constructive


   
Quote
(@henryf)
Reputable Member
Joined: 3 months ago
Posts: 291
 

You're right about the operational complexity. The cost and API rate limits are a huge factor.

I tried it for a FAQ section. The voice quality was excellent, but pre-rendering audio for dynamic content became a pipeline nightmare. We had to cache everything, and cache invalidation on content updates was messy.

For WCAG, check if their player offers full keyboard controls and ARIA labels. Last I looked, their embeddable player was more geared for listening posts than accessible UI components. You might be building a lot of custom controls to meet compliance, which defeats the purpose of using a managed service.



   
ReplyQuote
(@eval_engineer_101)
Reputable Member
Joined: 3 months ago
Posts: 283
 

That's a really interesting breakdown of the problem. I'm new to this specific use case, but your point about comprehension vs operational complexity is exactly what I'd be worried about.

Has anyone done a formal study comparing comprehension rates for a high-quality voice like WellSaid against, say, a high-end system voice like Microsoft Azure Neural? The cost difference is huge, but if the user experience and understanding are significantly better, maybe it's justified for key content.

Also, how does WellSaid's latency for real-time generation compare to other API-driven services? For dynamic content, even a half-second delay might feel awkward compared to a system voice that's instant.



   
ReplyQuote