MuseTalk is a local AI lip-sync model for creators and developers working on video dubbing or virtual avatars. It edits the face in an existing video to match supplied speech, including Chinese, English and Japanese audio. It runs on Windows and Linux with NVIDIA GPUs, and can process videos generated by MuseV.
You can use its Gradio browser interface locally; a separate hosted demo runs on Hugging Face Spaces. The interface includes a first-frame preview for checking facial results before generating a full video. MuseTalk accepts video, still images or image sequences alongside an audio file, and provides both batch generation and real-time avatar inference.
Hardware affects the use case. The project reports real-time generation on an NVIDIA Tesla V100. It also tested a GeForce RTX 3050 Ti laptop GPU with 4GB VRAM, where an eight-second video took about five minutes in reduced-precision mode. Training your own model requires substantially more GPU memory, and the project includes training code, data preparation tools and pretrained weights.
MuseTalk generates facial edits in a single pass through an encoded image representation. It uses Whisper-tiny to encode audio and a network adapted from Stable Diffusion, but doesn't use diffusion sampling. The code carries an MIT license, and the MuseTalk model weights permit commercial use; other included models retain their own licenses.
Results can include facial jitter or changes to mustaches, lip shape and lip color. The model edits a 256-by-256 face region, and GFPGAN can provide additional resolution enhancement.
Claim this page and we'll verify you by hand. MuseTalk gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find MuseTalk?Promote it
Something wrong or outdated on this page?
4.3KUpdated 6 months agoApache-2.0
Linux · Web#Hugging Face integration#Multimodal input
EchoMimic turns a portrait image and an audio recording into an animated talking-head video on your own hardware. Its distinguishing feature is editable facial landmark control: you can drive animation with audio, landmarks, or audio combined with selected landmarks. It's intended for academic research and suits people comparing methods for speech animation and facial motion control.
6.1KUpdated 1 year agoApache-2.0
Web#Batch processing#Hugging Face integration#Multimodal input
LatentSync is an open-source AI lip-sync tool that edits a video's mouth movements to match supplied audio. It runs on your own GPU and suits video creators working with talking faces or virtual avatars, as well as researchers who want to train their own lip-sync models. The code uses the Apache 2.0 license.
8.7KUpdated 2 years agoMIT
Linux#Hugging Face integration#Multimodal input#ONNX
Hallo turns a single portrait and a speech recording into an animated talking video on your own hardware. It's a local AI tool for creators working with talking portraits and researchers who want access to both generation and training code. The Python code uses the MIT license; required pretrained models and dependencies have their own terms.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
18.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.
17.3KUpdated 11 months agoApache-2.0
Windows · Linux · Web#Hugging Face integration#Multimodal input
FramePack is an open source desktop app for making videos from a still image and a written motion prompt. It runs on Windows and Linux, with generation handled by your own NVIDIA GPU. It suits people who want to make AI video locally and see the clip develop as it renders.