Favicon of EchoMimic

EchoMimic

Local AI portrait animation software turns images and audio into talking-head videos, with editable facial landmarks and Apache 2.0 source code.

EchoMimic turns a portrait image and an audio recording into an animated talking-head video on your own hardware. Its distinguishing feature is editable facial landmark control: you can drive animation with audio, landmarks, or audio combined with selected landmarks. It's intended for academic research and suits people comparing methods for speech animation and facial motion control.

The pretrained models were trained on English and Mandarin Chinese. Examples include spoken audio in both languages and singing. Pose control gives you another way to guide the animation, and a motion alignment tool matches movement between a reference portrait and a driving video. Accelerated models are available for both audio-only generation and combined audio and landmark control.

The Python project is open source under Apache 2.0. It runs locally, with tested Linux environments including CentOS and Ubuntu, and uses NVIDIA GPUs through CUDA. Tested hardware includes the V100 with 16 GB of GPU memory, RTX 4090D with 24 GB, and A100 with 80 GB; these are tested configurations rather than stated minimum requirements. Its model components include Stable Diffusion image variations and a VAE, plus Whisper for audio processing.

A local Gradio interface provides browser access, and a community ComfyUI integration is available. Hosted demos on Hugging Face and ModelScope provide a separate way to try the tool using remote GPUs.

Similar to EchoMimic