
Wan2.2 is an open source family of video generation models for creators and developers who want to make clips on their own GPUs. Licensed under Apache 2.0, it covers text-to-video and image-to-video generation, with separate models for speech-driven video and character animation. Generation runs on your hardware when you use the downloadable models.
The family includes dedicated text and image models, plus TI2V-5B, which accepts either kind of input. They support output up to 720p. The combined model uses compressed video representations to make that resolution practical on a single consumer GPU. The larger A14B models use a mixture-of-experts design: one expert shapes the scene early in generation, while another refines details later. Training data labeled for lighting, composition, contrast and color tone supports more control over a video's visual style.
Wan also operates a hosted creative platform powered by Alibaba Cloud. Its web service and mobile app provide image creation, video generation and editing, with API access available for the platform. Those services run in the cloud. The Wan2.2 model files are distributed through Hugging Face and ModelScope for local use.
Claim this page with an email at wan.video. Wan2.2 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Wan2.2?Promote it
Something wrong or outdated on this page?
135.6KUpdated 47 minutes agoGPL-3.0
macOS · Windows · Linux · Web#ControlNet#Inpainting#LoRA
ComfyUI is a local visual AI workspace for artists and technical teams who want to control how images, video, audio, 3D models and text are made. Its node canvas shows each model and processing step, so users can build and adjust workflows without writing code. It runs on your hardware.
12.3KUpdated 2 years agoApache-2.0
Web#Hugging Face integration#LoRA#Multimodal input
13KUpdated 11 months agoApache-2.0
Windows · Web#Hugging Face integration#LoRA#Multimodal input
34.6KUpdated 1 day agoApache-2.0
macOS#ControlNet#Hugging Face integration#Image-to-image
4.3KUpdated 6 months agoApache-2.0
Linux · Web#Hugging Face integration#Multimodal input
17.3KUpdated 11 months agoApache-2.0
Windows · Linux · Web#Hugging Face integration#Multimodal input
AnimateDiff adds text-driven animation to personalized Stable Diffusion models without requiring separate training for each model. It's for artists and developers who want to generate motion while keeping the visual style of a chosen image model. The Python implementation runs locally and includes a Gradio browser interface.
CogVideoX is a family of downloadable video generation models for developers, researchers and creators who want to generate clips on their own hardware. It turns English text prompts into video, animates a supplied image and can continue an existing video. A local Gradio web interface provides a browser front end for generation.
Diffusers is an open-source Python library for developers and researchers who want to run diffusion models on their own hardware or build generation features into an application. It uses PyTorch and supports image, video and audio generation. The library is licensed under Apache 2.0 and supports Apple Silicon.
EchoMimic turns a portrait image and an audio recording into an animated talking-head video on your own hardware. Its distinguishing feature is editable facial landmark control: you can drive animation with audio, landmarks, or audio combined with selected landmarks. It's intended for academic research and suits people comparing methods for speech animation and facial motion control.
FramePack is an open source desktop app for making videos from a still image and a written motion prompt. It runs on Windows and Linux, with generation handled by your own NVIDIA GPU. It suits people who want to make AI video locally and see the clip develop as it renders.