Wav2Lip is a local AI lip-sync tool that changes a face's mouth movements in an existing video to match supplied speech. It's for researchers and people making academic or personal video projects who want to process their own files. Commercial use is prohibited under the project's stated terms because its models were trained on the LRS2 dataset.
The model supports different speakers and languages, including CGI faces and synthetic voices. It accepts WAV and MP3 audio, as well as audio extracted from video files through FFmpeg. The output is a video with lip movements matched to the replacement audio.
Two pretrained models offer a concrete tradeoff: Wav2Lip prioritizes speech synchronization, while Wav2Lip + GAN improves visual quality at some cost to synchronization accuracy. Face framing and source resolution can affect the result. The models were trained on lower-resolution faces, so higher-resolution input doesn't always produce a better-looking output.
The Python code runs locally, and a Docker option is available. The project includes training code and evaluation benchmarks for researchers who want to train or compare models. Training on another dataset can require code changes and a separately trained lip-sync discriminator; fine-tuning on a few minutes of one speaker isn't a supported shortcut.
Sync Labs provides a separate hosted commercial offering. Its Sync API processes video and audio in the cloud and requires an account API key; Google Colab notebooks also provide a cloud route for the research model.
Claim this page and we'll verify you by hand. Wav2Lip gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Wav2Lip?Promote it
Something wrong or outdated on this page?
8.7KUpdated 2 years agoMIT
Linux#Hugging Face integration#Multimodal input#ONNX
Hallo turns a single portrait and a speech recording into an animated talking video on your own hardware. It's a local AI tool for creators working with talking portraits and researchers who want access to both generation and training code. The Python code uses the MIT license; required pretrained models and dependencies have their own terms.
6.1KUpdated 1 year agoApache-2.0
Web#Batch processing#Hugging Face integration#Multimodal input
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
2.4KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#Hugging Face integration
2.3KUpdated 1 day agoMPL-2.0
macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access
LatentSync is an open-source AI lip-sync tool that edits a video's mouth movements to match supplied audio. It runs on your own GPU and suits video creators working with talking faces or virtual avatars, as well as researchers who want to train their own lip-sync models. The code uses the Apache 2.0 license.
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
AllTalk TTS generates speech on your own computer. The project recommends v2 for most users; the saved documentation below describes v1, built on Coqui TTS and XTTSv2 models. It's for people adding voices to AI conversations or producing spoken audio from longer texts. It runs as a standalone application or alongside Text-generation-webui, with support for Windows, Linux and macOS.
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.