Skip to content
Notifications
Clear all

Did you see the new lip-sync beta? Does it look natural yet?

1 Posts
1 Users
0 Reactions
0 Views
(@jackson2m)
Estimable Member
Joined: 1 week ago
Posts: 67
Topic starter   [#13248]

Having spent the better part of the last quarter deep in vendor demos for AI-driven training modules, I've been rigorously testing the lip-sync capabilities across several platforms. Fliki's recent beta feature was, of course, immediately added to my evaluation matrix.

From my controlled tests using a standardized script (a 45-second technical explanation of SKU rationalization), I've observed the following:

**Strengths in the Current Beta:**
* Phoneme timing has improved markedly over the previous generation. Consonant-heavy phrases like "batch processing throughput" now show clear, distinct mouth movements for the plosives ('p', 'b', 't').
* The system handles mid-vowel transitions better than expected. The viseme mapping for diphthongs in words like "warehouse" or "cycle" is far less robotic.
* A noticeable reduction in the "uncanny valley" effect during neutral, declarative sentences—the area where most B2B explainer content lives.

**Persistent Shortcomings Noted:**
* Emotional cadence remains largely absent. A sentence about "critical supply chain disruption" is delivered with the same flat mouth shape as one about "routine inventory reconciliation." There is no tightening or emphasis.
* Coarticulation—how sounds blend in natural speech—is still underdeveloped. The phrase "financial software integration" runs together awkwardly, with the mouth failing to prepare for the 'sh' in 'software' during the preceding 'al'.
* Lateral movements and jaw drop on open vowels ('a' as in "asset") can appear exaggerated, particularly on certain male voice models. This becomes distracting in longer segments.

My current assessment is that for internal workflow automation tutorials or data-driven reports where absolute naturalism is secondary to clarity and speed, the beta is now a viable tool. However, for any customer-facing content where trust and nuanced communication are paramount—such as explaining a complex ERP migration or a sensitive billing policy—it does not yet pass the threshold. The lack of affective correlation between the audio tone (which can be modulated) and the visual output creates a subtle but perceptible dissonance.

I am curious to see comparative analysis from others using different voice models or longer-form content. Has anyone performed A/B testing with a human presenter against the Fliki output for viewer retention metrics on technical topics?


Data over opinions


   
Quote