Favicon of MuseTalk

MuseTalk

Local AI lip-sync model for Windows and Linux that matches faces to supplied audio, with NVIDIA GPU support and an MIT-licensed codebase.

MuseTalk is a local AI lip-sync model for creators and developers working on video dubbing or virtual avatars. It edits the face in an existing video to match supplied speech, including Chinese, English and Japanese audio. It runs on Windows and Linux with NVIDIA GPUs, and can process videos generated by MuseV.

You can use its Gradio browser interface locally; a separate hosted demo runs on Hugging Face Spaces. The interface includes a first-frame preview for checking facial results before generating a full video. MuseTalk accepts video, still images or image sequences alongside an audio file, and provides both batch generation and real-time avatar inference.

Hardware affects the use case. The project reports real-time generation on an NVIDIA Tesla V100. It also tested a GeForce RTX 3050 Ti laptop GPU with 4GB VRAM, where an eight-second video took about five minutes in reduced-precision mode. Training your own model requires substantially more GPU memory, and the project includes training code, data preparation tools and pretrained weights.

MuseTalk generates facial edits in a single pass through an encoded image representation. It uses Whisper-tiny to encode audio and a network adapted from Stable Diffusion, but doesn't use diffusion sampling. The code carries an MIT license, and the MuseTalk model weights permit commercial use; other included models retain their own licenses.

Results can include facial jitter or changes to mustaches, lip shape and lip color. The model edits a 256-by-256 face region, and GFPGAN can provide additional resolution enhancement.

Similar to MuseTalk