
HunyuanVideo is an AI video generation model for creators and developers who want to generate footage on their own hardware. Tencent provides model weights and inference code for text-to-video and image-to-video generation, alongside a hosted web experience. Local inference runs on your GPUs; the web offering runs through Tencent's service.
The model can produce realistic or stylized scenes, with an emphasis on continuous movement and sequences of actions within a shot. It also supports camera movements and cuts between views. These capabilities suit advertising concepts, film experiments and animated scenes where motion and framing matter as much as the appearance of individual frames.
For local workflows, HunyuanVideo works with ComfyUI and Hugging Face Diffusers, and includes a Gradio web demo. FP8 quantized weights provide an alternative to the full weights, while multi-GPU inference can use several GPUs to speed up generation. The repository includes PyTorch model definitions for developers who want to work directly with the model.
A separate Hunyuan-Large prompt rewrite model can expand descriptions of composition, lighting and camera movement. Its more visually focused mode can lose some details of the original request. Underneath, HunyuanVideo combines a diffusion transformer with a multimodal language model that interprets prompts, and uses a shared architecture for image and video generation. Tencent also provides the prompt rewrite model's weights for local deployment.
Claim this page with an email at aivideo.hunyuan.tencent.com. HunyuanVideo gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find HunyuanVideo?Promote it
Something wrong or outdated on this page?
13KUpdated 11 months agoApache-2.0
Windows · Web#Hugging Face integration#LoRA#Multimodal input
CogVideoX is a family of downloadable video generation models for developers, researchers and creators who want to generate clips on their own hardware. It turns English text prompts into video, animates a supplied image and can continue an existing video. A local Gradio web interface provides a browser front end for generation.
12.3KUpdated 2 years agoApache-2.0
Web#Hugging Face integration#LoRA#Multimodal input
17.3KUpdated 11 months agoApache-2.0
Windows · Linux · Web#Hugging Face integration#Multimodal input
3.2KUpdated 2 years agoMIT
macOS · Web#Hugging Face integration#Multimodal input
7.6KUpdated 8 months ago
Web#Hugging Face integration#Multimodal input
3.7KUpdated 11 months agoApache-2.0
Web#Hugging Face integration#LoRA
Mochi 1 is a text-to-video model for creators and developers who want to generate videos on their own hardware or adapt a model to their own footage. Genmo releases it under Apache 2.0, with downloadable weights and code for local use. Genmo also offers a hosted playground for trying the model in a browser.
AnimateDiff adds text-driven animation to personalized Stable Diffusion models without requiring separate training for each model. It's for artists and developers who want to generate motion while keeping the visual style of a chosen image model. The Python implementation runs locally and includes a Gradio browser interface.
FramePack is an open source desktop app for making videos from a still image and a written motion prompt. It runs on Windows and Linux, with generation handled by your own NVIDIA GPU. It suits people who want to make AI video locally and see the clip develop as it renders.
Pyramid Flow is an open-source AI video generator for people who want to create clips on their own hardware, and for researchers working on video models. It turns text prompts into video or animates a supplied image with guidance from text. The Python code uses PyTorch and carries MIT. Model licenses are separate: the SD3-derived weights use the Stability AI Community License. Check the chosen checkpoint before deployment.
SkyReels-V2 is a family of downloadable video models for developers and researchers who want to run generation on their own GPUs. It supports text-to-video, image-to-video and extended video generation. Its autoregressive Diffusion Forcing architecture generates longer sequences by continuing video segments.