LatentSync is an open-source AI lip-sync tool that edits a video's mouth movements to match supplied audio. It runs on your own GPU and suits video creators working with talking faces or virtual avatars, as well as researchers who want to train their own lip-sync models. The code uses the Apache 2.0 license.
Its approach builds on Stable Diffusion and AnimateDiff. Rather than generating a separate motion representation first, it uses audio to guide video generation directly. Whisper supplies the audio features, while temporal layers help keep movement consistent between frames. The model trains on higher-resolution video to reduce blur, and the demos cover both filmed faces and anime characters.
A Gradio browser interface provides a visual way to process videos, and a command-line interface supports scripted work. Model checkpoints are available through Hugging Face. The higher-resolution model needs at least 18 GB of GPU memory for inference. Generation settings let users trade processing speed against visual quality, though stronger audio guidance can introduce distortion or jitter.
The project includes training code and a data preparation pipeline for people working with their own footage. That pipeline splits scenes, aligns faces, corrects audio timing and filters clips for visual quality and synchronization. An efficient training setup fits in 20 GB of VRAM, with some loss of visual quality and movement consistency. SyncNet evaluation tools measure synchronization confidence in generated videos and assess model accuracy on datasets.
Claim this page and we'll verify you by hand. LatentSync gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find LatentSync?Promote it
Something wrong or outdated on this page?
8.7KUpdated 2 years agoMIT
Linux#Hugging Face integration#Multimodal input#ONNX
Hallo turns a single portrait and a speech recording into an animated talking video on your own hardware. It's a local AI tool for creators working with talking portraits and researchers who want access to both generation and training code. The Python code uses the MIT license; required pretrained models and dependencies have their own terms.
4.3KUpdated 6 months agoApache-2.0
Linux · Web#Hugging Face integration#Multimodal input
6.6KUpdated 1 year ago
Windows · Linux · Web#Batch processing#Inpainting#Multilingual
MuseTalk is a local AI lip-sync model for creators and developers working on video dubbing or virtual avatars. It edits the face in an existing video to match supplied speech, including Chinese, English and Japanese audio. It runs on Windows and Linux with NVIDIA GPUs, and can process videos generated by MuseV.
3.2KUpdated 2 years agoMIT
macOS · Web#Hugging Face integration#Multimodal input
29.9KUpdated 6 months agoApache-2.0
#Hugging Face integration#Multimodal input
Open-Sora is an open source AI video generation project for developers, researchers, and creators who want to run and adapt a model on their own hardware. Its model focuses on turning reference images into video, with text prompts guiding the result. It also generates video directly from text. The code and Open-Sora 2.0 weights use Apache 2.0.
13.2KUpdated 1 year ago
Linux · Docker#Multilingual#Multimodal input
Wav2Lip is a local AI lip-sync tool that changes a face's mouth movements in an existing video to match supplied speech. It's for researchers and people making academic or personal video projects who want to process their own files. Commercial use is prohibited under the project's stated terms because its models were trained on the LRS2 dataset.
EchoMimic turns a portrait image and an audio recording into an animated talking-head video on your own hardware. Its distinguishing feature is editable facial landmark control: you can drive animation with audio, landmarks, or audio combined with selected landmarks. It's intended for academic research and suits people comparing methods for speech animation and facial motion control.
Pyramid Flow is an open-source AI video generator for people who want to create clips on their own hardware, and for researchers working on video models. It turns text prompts into video or animates a supplied image with guidance from text. The Python code uses PyTorch and carries MIT. Model licenses are separate: the SD3-derived weights use the Stability AI Community License. Check the chosen checkpoint before deployment.