I've been researching AI voice and video localization tools for our customer onboarding content. We have a library of YouTube tutorials that need to be available in multiple languages, and manual dubbing is off the table budget-wise.
I've read through Resemble's documentation and some general reviews, but I'm struggling to find concrete, user-reported details on the video dubbing workflow. Specifically, the lip-sync aspect. Their site mentions "proprietary lip-syncing technology," but that's quite vague.
For anyone who has actually run a project through them:
* How does the lip-sync hold up for a typical talking-head YouTube video? Is it just a basic audio alignment, or does it attempt to modify the video track?
* What's the actual process like? Do you just upload a video and a translated script, or is there a manual review/alignment step required?
* Most importantly, does the output feel "uncanny" or passable for professional B2B content where clarity is more important than perfect realism?
I'm comparing this against a few other platforms, but Resemble's voice cloning features are interesting for maintaining a consistent brand voice across languages. The lip-sync potential is the main unknown for me.