Hallo turns a single portrait and a speech recording into an animated talking video on your own hardware. It's a local AI tool for creators working with talking portraits and researchers who want access to both generation and training code. The Python code uses the MIT license; required pretrained models and dependencies have their own terms.
The animation follows the supplied audio, with separate controls for head pose, facial movement and lips. Hallo builds on Stable Diffusion 1.5 and AnimateDiff, and uses Wav2Vec to process speech. It exports MP4 video.
Input suitability matters. Hallo expects a square portrait with a prominent, mostly forward-facing face, so side profiles aren't a good fit. Its speech input is English WAV audio, reflecting the language of its training data. Vocals should be clear, though recordings can include background music.
The documented platform is Ubuntu with CUDA, and testing used an NVIDIA A100 GPU. Pretrained models are available through Hugging Face for local generation. Community projects provide a Windows adaptation, a browser interface, ComfyUI integration and a Docker image. A hosted Gradio demo on Hugging Face offers a separate way to try it.
For researchers adapting the model, Hallo includes data preparation and training code for talking-face video datasets. Training supports multiple GPUs and multiple machines through Hugging Face Accelerate. A separate community extension, JoyHallo, adds Mandarin support.
Claim this page and we'll verify you by hand. Hallo gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Hallo?Promote it
Something wrong or outdated on this page?
6.1KUpdated 1 year agoApache-2.0
Web#Batch processing#Hugging Face integration#Multimodal input
LatentSync is an open-source AI lip-sync tool that edits a video's mouth movements to match supplied audio. It runs on your own GPU and suits video creators working with talking faces or virtual avatars, as well as researchers who want to train their own lip-sync models. The code uses the Apache 2.0 license.
4.3KUpdated 6 months agoApache-2.0
Linux · Web#Hugging Face integration#Multimodal input
6.6KUpdated 1 year ago
Windows · Linux · Web#Batch processing#Inpainting#Multilingual
MuseTalk is a local AI lip-sync model for creators and developers working on video dubbing or virtual avatars. It edits the face in an existing video to match supplied speech, including Chinese, English and Japanese audio. It runs on Windows and Linux with NVIDIA GPUs, and can process videos generated by MuseV.
13.2KUpdated 1 year ago
Linux · Docker#Multilingual#Multimodal input
Wav2Lip is a local AI lip-sync tool that changes a face's mouth movements in an existing video to match supplied speech. It's for researchers and people making academic or personal video projects who want to process their own files. Commercial use is prohibited under the project's stated terms because its models were trained on the LRS2 dataset.
29.9KUpdated 6 months agoApache-2.0
#Hugging Face integration#Multimodal input
Open-Sora is an open source AI video generation project for developers, researchers, and creators who want to run and adapt a model on their own hardware. Its model focuses on turning reference images into video, with text prompts guiding the result. It also generates video directly from text. The code and Open-Sora 2.0 weights use Apache 2.0.
3.2KUpdated 2 years agoMIT
macOS · Web#Hugging Face integration#Multimodal input
EchoMimic turns a portrait image and an audio recording into an animated talking-head video on your own hardware. Its distinguishing feature is editable facial landmark control: you can drive animation with audio, landmarks, or audio combined with selected landmarks. It's intended for academic research and suits people comparing methods for speech animation and facial motion control.
Pyramid Flow is an open-source AI video generator for people who want to create clips on their own hardware, and for researchers working on video models. It turns text prompts into video or animates a supplied image with guidance from text. The Python code uses PyTorch and carries MIT. Model licenses are separate: the SD3-derived weights use the Stability AI Community License. Check the chosen checkpoint before deployment.