Favicon of Hallo

Hallo

An open-source portrait animation tool that runs on Ubuntu with an NVIDIA GPU, turning a still image and English speech into a talking video.

Hallo turns a single portrait and a speech recording into an animated talking video on your own hardware. It's a local AI tool for creators working with talking portraits and researchers who want access to both generation and training code. The Python code uses the MIT license; required pretrained models and dependencies have their own terms.

The animation follows the supplied audio, with separate controls for head pose, facial movement and lips. Hallo builds on Stable Diffusion 1.5 and AnimateDiff, and uses Wav2Vec to process speech. It exports MP4 video.

Input suitability matters. Hallo expects a square portrait with a prominent, mostly forward-facing face, so side profiles aren't a good fit. Its speech input is English WAV audio, reflecting the language of its training data. Vocals should be clear, though recordings can include background music.

The documented platform is Ubuntu with CUDA, and testing used an NVIDIA A100 GPU. Pretrained models are available through Hugging Face for local generation. Community projects provide a Windows adaptation, a browser interface, ComfyUI integration and a Docker image. A hosted Gradio demo on Hugging Face offers a separate way to try it.

For researchers adapting the model, Hallo includes data preparation and training code for talking-face video datasets. Training supports multiple GPUs and multiple machines through Hugging Face Accelerate. A separate community extension, JoyHallo, adds Mandarin support.

Similar to Hallo