Favicon of HunyuanVideo

HunyuanVideo

AI video generation model you can run on your own GPUs, with text-to-video and image-to-video support, ComfyUI integration and downloadable weights.

Screenshot of HunyuanVideo website

HunyuanVideo is an AI video generation model for creators and developers who want to generate footage on their own hardware. Tencent provides model weights and inference code for text-to-video and image-to-video generation, alongside a hosted web experience. Local inference runs on your GPUs; the web offering runs through Tencent's service.

The model can produce realistic or stylized scenes, with an emphasis on continuous movement and sequences of actions within a shot. It also supports camera movements and cuts between views. These capabilities suit advertising concepts, film experiments and animated scenes where motion and framing matter as much as the appearance of individual frames.

For local workflows, HunyuanVideo works with ComfyUI and Hugging Face Diffusers, and includes a Gradio web demo. FP8 quantized weights provide an alternative to the full weights, while multi-GPU inference can use several GPUs to speed up generation. The repository includes PyTorch model definitions for developers who want to work directly with the model.

A separate Hunyuan-Large prompt rewrite model can expand descriptions of composition, lighting and camera movement. Its more visually focused mode can lose some details of the original request. Underneath, HunyuanVideo combines a diffusion transformer with a multimodal language model that interprets prompts, and uses a shared architecture for image and video generation. Tencent also provides the prompt rewrite model's weights for local deployment.

Similar to HunyuanVideo