EchoMimic turns a portrait image and an audio recording into an animated talking-head video on your own hardware. Its distinguishing feature is editable facial landmark control: you can drive animation with audio, landmarks, or audio combined with selected landmarks. It's intended for academic research and suits people comparing methods for speech animation and facial motion control.
The pretrained models were trained on English and Mandarin Chinese. Examples include spoken audio in both languages and singing. Pose control gives you another way to guide the animation, and a motion alignment tool matches movement between a reference portrait and a driving video. Accelerated models are available for both audio-only generation and combined audio and landmark control.
The Python project is open source under Apache 2.0. It runs locally, with tested Linux environments including CentOS and Ubuntu, and uses NVIDIA GPUs through CUDA. Tested hardware includes the V100 with 16 GB of GPU memory, RTX 4090D with 24 GB, and A100 with 80 GB; these are tested configurations rather than stated minimum requirements. Its model components include Stable Diffusion image variations and a VAE, plus Whisper for audio processing.
A local Gradio interface provides browser access, and a community ComfyUI integration is available. Hosted demos on Hugging Face and ModelScope provide a separate way to try the tool using remote GPUs.
Claim this page and we'll verify you by hand. EchoMimic gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find EchoMimic?Promote it
Something wrong or outdated on this page?
6.6KUpdated 1 year ago
Windows · Linux · Web#Batch processing#Inpainting#Multilingual
MuseTalk is a local AI lip-sync model for creators and developers working on video dubbing or virtual avatars. It edits the face in an existing video to match supplied speech, including Chinese, English and Japanese audio. It runs on Windows and Linux with NVIDIA GPUs, and can process videos generated by MuseV.
8.7KUpdated 2 years agoMIT
Linux#Hugging Face integration#Multimodal input#ONNX
6.1KUpdated 1 year agoApache-2.0
Web#Batch processing#Hugging Face integration#Multimodal input
17.3KUpdated 11 months agoApache-2.0
Windows · Linux · Web#Hugging Face integration#Multimodal input
19.1KUpdated 4 months ago
macOS · Windows · Linux · Web#Hugging Face integration
LivePortrait animates a still portrait using facial expressions and head movements from a driving video. It's for creators who want to animate faces or edit motion in portrait videos on their own hardware. It works with realistic photos as well as portraits in oil paintings, sculptures and 3D renders.
14.1KUpdated 3 years ago
macOS · Windows · Linux · Web#Multimodal input#Works offline
SadTalker generates talking head videos from a single portrait and an audio recording. The project states an Apache 2.0 license and removal of its earlier noncommercial restriction. It runs locally on Windows, Linux and macOS, and suits creators who want to animate a face without recording a person on camera. Its animation includes facial expressions and head movement, with examples covering speech and singing in different languages.
Hallo turns a single portrait and a speech recording into an animated talking video on your own hardware. It's a local AI tool for creators working with talking portraits and researchers who want access to both generation and training code. The Python code uses the MIT license; required pretrained models and dependencies have their own terms.
LatentSync is an open-source AI lip-sync tool that edits a video's mouth movements to match supplied audio. It runs on your own GPU and suits video creators working with talking faces or virtual avatars, as well as researchers who want to train their own lip-sync models. The code uses the Apache 2.0 license.
FramePack is an open source desktop app for making videos from a still image and a written motion prompt. It runs on Windows and Linux, with generation handled by your own NVIDIA GPU. It suits people who want to make AI video locally and see the clip develop as it renders.