Favicon of LatentSync

LatentSync

Open-source local AI lip-sync tool with a Gradio interface and Apache 2.0 code license. GPU inference requires 8 GB or 18 GB VRAM, depending on the model.

LatentSync is an open-source AI lip-sync tool that edits a video's mouth movements to match supplied audio. It runs on your own GPU and suits video creators working with talking faces or virtual avatars, as well as researchers who want to train their own lip-sync models. The code uses the Apache 2.0 license.

Its approach builds on Stable Diffusion and AnimateDiff. Rather than generating a separate motion representation first, it uses audio to guide video generation directly. Whisper supplies the audio features, while temporal layers help keep movement consistent between frames. The model trains on higher-resolution video to reduce blur, and the demos cover both filmed faces and anime characters.

A Gradio browser interface provides a visual way to process videos, and a command-line interface supports scripted work. Model checkpoints are available through Hugging Face. The higher-resolution model needs at least 18 GB of GPU memory for inference. Generation settings let users trade processing speed against visual quality, though stronger audio guidance can introduce distortion or jitter.

The project includes training code and a data preparation pipeline for people working with their own footage. That pipeline splits scenes, aligns faces, corrects audio timing and filters clips for visual quality and synchronization. An efficient training setup fits in 20 GB of VRAM, with some loss of visual quality and movement consistency. SyncNet evaluation tools measure synchronization confidence in generated videos and assess model accuracy on datasets.

Similar to LatentSync