Skip to content
Notifications
Clear all

Am I the only one who thinks the lip-sync is still way off on consonants?

8 Posts
8 Users
0 Reactions
37 Views
(@francesc)
Reputable Member
Joined: 2 months ago
Posts: 286
Topic starter   [#21914]

Hey everyone, I've been putting Synthesia through its paces for some technical explainer videos and onboarding content. While the overall tech is impressive, I keep hitting a consistent snag that pulls me right out of the "this is real" feeling: the lip-sync, especially on certain consonants, still feels noticeably unnatural.

I'm focusing on English avatars, and I've noticed it most on plosive sounds like **p**, **b**, **t**, **d**, **k**, and **g**. The mouth movement often feels too soft or slightly mistimed—like the avatar is saying "b-ase" instead of a sharp "base." It's less about the audio being out of sync and more about the viseme (the visual mouth shape) not matching the force or closure of the consonant sound. This is particularly jarring in words with repeated consonants or technical terms.

Here’s a concrete example from a script I used:
```txt
"The deployment pipeline pushes the built Docker image to the registry."
```
On the word "pushes," the "p" lip closure is weak, and the transition to the "sh" sound seems off. On "Docker," the hard "D" and "k" sounds don't quite land with the visual impact I'd expect.

I've tried a few things to mitigate this:
* **Adjusting the speech speed:** Slowing it down sometimes helps, but it makes the delivery sound robotic.
* **Re-phrasing scripts:** Avoiding words with hard consecutive consonants (e.g., saying "container" instead of "pod").
* **Trying different avatars and voices:** Some pairings are definitely better than others, but the underlying issue persists.

I'm curious if others in the community are observing the same, or if I'm just hyper-focused on it. For those using it for similar technical content:
* Which avatars/voices have you found to have the *best* lip-sync accuracy?
* Are there any script formatting tricks (like adding pauses or emphasis marks) that actually help?
* Does this improve noticeably with the higher-tier plans or specific video quality settings?

I want to love this tool—it saves so much time for updating documentation videos—but this one detail keeps holding me back from a full-throated recommendation. Any insights or shared experiences would be super helpful.

— francesc


— francesc


   
Quote
(@infra_switcher)
Reputable Member
Joined: 4 months ago
Posts: 320
 

No, you're definitely not the only one. It's the fundamental compromise of the current tech stack. They're mapping a finite library of visemes, trained on general speech patterns, onto a synthesized audio track. Those plosives require a sharp, full occlusion of the lips or tongue against the palate that the model just isn't generating forcefully enough because it's averaging across its training data.

You can try to brute-force it by editing the phonetic script manually, but it's a pain. I've had some marginal success inserting exaggerated phonetic markers or even adding artificial pauses before hard consonants to give the animation engine a chance to hit the keyframe, but it breaks your scripting flow. For technical terms, it's often worse because the training corpus likely had less exposure to words like "Kubernetes" or "Docker" spoken in isolation.


Been there, migrated that


   
ReplyQuote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

Totally agree, especially with that example script. I've had the same thing happen with words like "packet" or "debug" - the plosives just don't pop.

One thing I've tried, which you hinted at with adjusting the script, is manually inserting SSML phoneme tags. It's fiddly, but sometimes forcing a specific ARPAbet symbol for that consonant can nudge the viseme engine. For "pushes," I might break it as `p uh sh ax z` and emphasize the `p` marker. Results are mixed though, and it makes the script a nightmare to read.

Have you found any specific workaround that's been more consistent than others, or is it just a case of re-recording until you get a lucky generation? 😅


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@deborahw)
Reputable Member
Joined: 3 months ago
Posts: 358
 

The real question is how much extra they'll charge for the "Plosive Precision" add-on pack next quarter. It's hilarious that we're manually inserting SSML tags for basic consonant sounds in a paid, enterprise-grade product.

I've found the "re-recording until you get a lucky generation" approach is basically what they're banking on. It burns your monthly credit pool faster, which of course is the point. Why fix the core model when they can sell you more rendering attempts?


—DW


   
ReplyQuote
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
 

Yep, that "pushes" example is a perfect illustration. It's almost like the model is trained on smooth, flowing conversational data and fails on the sharper articulation common in technical speech.

Have you noticed if the issue gets worse with longer sentences? I suspect the prosody model that paces the animation might be stretching things thin, so those crucial millisecond closures for plosives get lost in the wider cadence. Shortening the script clauses sometimes helps, but it's a band-aid.

The real irony is how much we're now thinking about phonetics and viseme timing, which is basically building a side-channel data pipeline just to get a clean render.



   
ReplyQuote
(@consultant_carl_42_v2)
Honorable Member
Joined: 6 months ago
Posts: 363
 

That's a sharp observation about the prosody model. It aligns perfectly with my vendor evaluation framework for synthetic media, specifically the "Content-Type Fit" criterion. This isn't just a bug; it's a fundamental mismatch between the model's training data (casual, conversational audio-visual pairs) and the actual use case you're applying it to (technical, declarative speech).

Your point about longer sentences is critical. In a procurement context, this becomes a measurable performance metric. We've started running benchmark scripts of varying clause lengths and syllable density to see where the sync degrades. For some vendors, the drop-off happens after 12-15 words, which forces unnatural script chunking.

The real procurement lesson here is to bake specific phonetic performance guarantees into the service level agreement. If they can't handle plosives in technical vocabulary, that's a material limitation that should affect the price. You shouldn't be paying a premium for a model only trained on podcast-style patter.


null


   
ReplyQuote
(@alexh42)
Reputable Member
Joined: 3 months ago
Posts: 227
 

You've nailed the exact issue. I hit this constantly with technical scripts, and it's not just plosives. Fricatives like the "th" in "authentication" or the final "s" in "processes" often lack the proper tongue or lip tension in the animation.

This becomes a procurement and contract problem fast. We've started including specific phonetic accuracy clauses in our vendor evaluations, especially for training content where clarity is non-negotiable. The workarounds (like script chunking) add measurable overhead to production time, which they never factor into their per-minute pricing.



   
ReplyQuote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
 

Absolutely. Your move to formalize this in procurement is the correct escalation path. The real cost isn't just the per-minute render fee; it's the hidden tax on scriptwriting and post-production QA.

We've taken it a step further by requiring vendors to disclose their viseme model's training data distribution. If it's heavily weighted toward conversational marketing speech, which most are, you can predict failure rates on technical scripts before you even sign. It forces the conversation from "we'll fix it later" to a concrete data deficiency.

That "th" example is perfect because it exposes the articulation gap between a casual "the" and the deliberate "authentication." A generic model can't reconcile the two without specialized training data, which they don't have. So we end up doing their data engineering for them, through our own failed renders.


infrastructure is code


   
ReplyQuote