Hey folks, hoping someone here has run into this and found a workaround. I've been trying to build a custom avatar in HeyGen for a demo project, but the lip sync on the generated videos is consistently off by a noticeable fraction of a second. It's just enough to feel uncanny and makes the output unusable.
I've tried:
* Multiple different audio tracks (clean studio recordings, different bitrates)
* Both the web upload and the API for generation
* Adjusting the avatar's "expressiveness" setting up and down
Reached out to support with a clear example, timestamps, and my workflow. Their response was basically "it looks fine on our end" and a suggestion to re-upload the source photo (which didn't help). Feels like a classic case of something being lost between the audio preprocessing and the model inference.
Has anyone else dealt with this? Specifically for custom avatars, not the stock ones. I'm wondering if there's a particular audio format or preprocessing step (maybe normalizing to a specific LUFS?) that gives the pipeline better input to work with. Or is this just a known limitation with the current model?
—Claire
The audio format angle is interesting. It's unlikely to be a bitrate issue if you've tried multiple clean sources. Have you checked the actual phoneme timing in your input tracks? Most of these systems don't work on raw audio, they run speech-to-text first and then map phonemes to visemes. If the STT step introduces a latency or misaligns the word boundaries, the lip sync will be permanently off from that point.
Try generating a transcript from your audio using HeyGen's own system if it offers one, or a different service, and compare the word timestamps. I've seen cases where background noise, even at low levels, or certain vocal tones can cause the speech recognition to incorrectly place a word boundary, shifting everything that follows. It's a preprocessing bug that's invisible unless you look at the intermediate data.
Also, did you test with an artificially simple audio track? Something like a slow, deliberate count from one to ten with clear pauses. That would rule out any linguistic complexity and isolate the sync issue to the pipeline's base latency.
—davidr
That "looks fine on our end" support response is the modern equivalent of a shrug. Classic black box vendor deflection.
The phoneme mapping theory from the other comment is a good angle, but I'd suspect the preprocessing pipeline itself is the culprit, especially for custom avatars. These systems often have separate optimization queues for their stock models versus the less-tuned, computationally heavier custom model inference. If there's any buffering or batch processing happening between the audio analysis and the video render, that's where your consistent fractional delay is getting baked in.
A crude but sometimes effective test: try generating the same audio with one of their default, non-custom avatars. If the sync is perfect there, you've isolated it to the custom model pipeline, and your only real leverage is to badger support with that comparative evidence. They can't claim it's your audio if it works elsewhere in their own system.
keep it simple
Been down this road with their custom avatars too. That "looks fine on our end" response is frustrating because it *is* likely a pipeline issue specific to custom models, like the other comments hint at.
Since you've ruled out audio source, try this workaround: insert a 100-200ms silence pad at the very start of your audio track before uploading. Not elegant, but if the delay is consistent, it might compensate for that baked-in processing lag and bring the visible lip movement forward enough to sync. I've had this trick work with other services.
Also, double-check your source photo. If the avatar's starting mouth position in the photo is slightly open or closed, it can throw off the initial viseme prediction for the first syllable, making the whole thing feel off. A neutral, relaxed mouth gave me better results. Good luck!
Data doesn't lie, but dashboards sometimes do.
Good point about phoneme mapping, but I think you're giving the STT system too much credit. If it's a commercial SaaS tool, the phoneme-to-viseme mapping is likely a pre-trained model bundle, not a live pipeline. The delay is probably baked in before the STT even runs.
Testing with a simple count won't reveal a processing lag if it's constant across the entire clip. You'd get the same offset on "one" as you would on "ten." The real test is whether the offset is *consistent*. If it's always, say, 120ms off, that's a pipeline latency issue. If it drifts, *then* you're into STT word-boundary problems.
Ever notice this only happens with the higher-res custom avatars? Their stock ones are probably on a more optimized, lower-latency inference path.
Interesting hack with the silence pad. If the latency is fixed and baked into the pipeline, that could actually work. Makes me wonder if the delay is in the initial frame buffer for the custom model render.
Has anyone noticed if the delay is exactly the same length every time? If it is, that silence trick might be a viable band-aid while they fix their pipeline.