LatentSync 1.5: Windows installation and lip sync tutorial

Learn to install LatentSync 1.5 on Windows, set up PyTorch and model files, and process matching video and audio with a stated 6–8 GB VRAM requirement.

Player not loading? Watch on YouTube

This tutorial covers a local AI lip sync setup using an older LatentSync 1.5 fork on Windows. The presenter states that it needs 6 to 8 GB of VRAM and chooses it over version 1.6, which they say requires more than 16 GB. Their preference for 1.5's output is a personal assessment, rather than a measured comparison.

The listed prerequisites are CUDA Toolkit 11.8 to 12.8, Python 3.10.11, Visual Studio, FFmpeg and Git for Windows. The walkthrough clones the repository into a chosen folder, creates and activates a Python virtual environment, and installs PyTorch separately. It then edits requirements.txt before installing the remaining dependencies. Those edits include removing torch and torchaudio entries, changing several package versions and removing another package. The presenter does not read out the replacement version numbers.

Next, the presenter creates a checkpoints folder, downloads model files and launches the browser interface. A run.bat file provides a shortcut for later launches. To generate a lip-synced video, the demonstration uploads a video and a separate audio file; the presenter says both should have the same duration. The interface also exposes scale, step count and seed settings for trying different results.