Favicon of Wav2Lip

Wav2Lip

Local AI lip-sync software matches video to new speech, with pretrained models and training code for research, academic and personal use.

Wav2Lip is a local AI lip-sync tool that changes a face's mouth movements in an existing video to match supplied speech. It's for researchers and people making academic or personal video projects who want to process their own files. Commercial use is prohibited under the project's stated terms because its models were trained on the LRS2 dataset.

The model supports different speakers and languages, including CGI faces and synthetic voices. It accepts WAV and MP3 audio, as well as audio extracted from video files through FFmpeg. The output is a video with lip movements matched to the replacement audio.

Two pretrained models offer a concrete tradeoff: Wav2Lip prioritizes speech synchronization, while Wav2Lip + GAN improves visual quality at some cost to synchronization accuracy. Face framing and source resolution can affect the result. The models were trained on lower-resolution faces, so higher-resolution input doesn't always produce a better-looking output.

The Python code runs locally, and a Docker option is available. The project includes training code and evaluation benchmarks for researchers who want to train or compare models. Training on another dataset can require code changes and a separately trained lip-sync discriminator; fine-tuning on a few minutes of one speaker isn't a supported shortcut.

Sync Labs provides a separate hosted commercial offering. Its Sync API processes video and audio in the cloud and requires an account API key; Google Colab notebooks also provide a cloud route for the research model.

Similar to Wav2Lip