17.7KUpdated 1 week agoApache-2.0
#Hugging Face integration
Wan2.2 is an open source family of video generation models for creators and developers who want to make clips on their own GPUs. Licensed under Apache 2.0, it covers text-to-video and image-to-video generation, with separate models for speech-driven video and character animation. Generation runs on your hardware when you use the downloadable models.
3.2KUpdated 2 years agoMIT
macOS · Web#Hugging Face integration#Multimodal input
Pyramid Flow is an open-source AI video generator for people who want to create clips on their own hardware, and for researchers working on video models. It turns text prompts into video or animates a supplied image with guidance from text. The Python code uses PyTorch and carries MIT. Model licenses are separate: the SD3-derived weights use the Stability AI Community License. Check the chosen checkpoint before deployment.
12.3KUpdated 2 years agoApache-2.0
Web#Hugging Face integration#LoRA#Multimodal input
AnimateDiff adds text-driven animation to personalized Stable Diffusion models without requiring separate training for each model. It's for artists and developers who want to generate motion while keeping the visual style of a chosen image model. The Python implementation runs locally and includes a Gradio browser interface.
27.3KUpdated 9 months agoMIT
Web#Hugging Face integration
Stable Video Diffusion is an image-to-video model from Stability AI that animates a supplied still image into a short clip. It is for developers and researchers working with video generation on their own hardware. The SVD model generates 14 frames; SVD-XT is fine-tuned for 25 frames at 576×1024 resolution.
29.9KUpdated 6 months agoApache-2.0
#Hugging Face integration#Multimodal input
Open-Sora is an open source AI video generation project for developers, researchers, and creators who want to run and adapt a model on their own hardware. Its model focuses on turning reference images into video, with text prompts guiding the result. It also generates video directly from text. The code and Open-Sora 2.0 weights use Apache 2.0.
7.6KUpdated 8 months ago
Web#Hugging Face integration#Multimodal input
SkyReels-V2 is a family of downloadable video models for developers and researchers who want to run generation on their own GPUs. It supports text-to-video, image-to-video and extended video generation. Its autoregressive Diffusion Forcing architecture generates longer sequences by continuing video segments.
3.7KUpdated 11 months agoApache-2.0
Web#Hugging Face integration#LoRA
Mochi 1 is a text-to-video model for creators and developers who want to generate videos on their own hardware or adapt a model to their own footage. Genmo releases it under Apache 2.0, with downloadable weights and code for local use. Genmo also offers a hosted playground for trying the model in a browser.
13KUpdated 11 months agoApache-2.0
Windows · Web#Hugging Face integration#LoRA#Multimodal input
CogVideoX is a family of downloadable video generation models for developers, researchers and creators who want to generate clips on their own hardware. It turns English text prompts into video, animates a supplied image and can continue an existing video. A local Gradio web interface provides a browser front end for generation.
8.7KUpdated 2 years agoMIT
Linux#Hugging Face integration#Multimodal input#ONNX
Hallo turns a single portrait and a speech recording into an animated talking video on your own hardware. It's a local AI tool for creators working with talking portraits and researchers who want access to both generation and training code. The Python code uses the MIT license; required pretrained models and dependencies have their own terms.
9.2KUpdated 2 weeks agoApache-2.0
#ControlNet#LoRA#Multimodal input
Sana is an open-source framework for running image and video generation on your own hardware, with image models small enough for laptop GPUs. It's aimed at creators who want local AI generation and developers who need training and inference pipelines for their own models. The code uses the Apache 2.0 license.
4.3KUpdated 6 months agoApache-2.0
Linux · Web#Hugging Face integration#Multimodal input
EchoMimic turns a portrait image and an audio recording into an animated talking-head video on your own hardware. Its distinguishing feature is editable facial landmark control: you can drive animation with audio, landmarks, or audio combined with selected landmarks. It's intended for academic research and suits people comparing methods for speech animation and facial motion control.
6.6KUpdated 1 year ago
Windows · Linux · Web#Batch processing#Inpainting#Multilingual
MuseTalk is a local AI lip-sync model for creators and developers working on video dubbing or virtual avatars. It edits the face in an existing video to match supplied speech, including Chinese, English and Japanese audio. It runs on Windows and Linux with NVIDIA GPUs, and can process videos generated by MuseV.
17.3KUpdated 11 months agoApache-2.0
Windows · Linux · Web#Hugging Face integration#Multimodal input
FramePack is an open source desktop app for making videos from a still image and a written motion prompt. It runs on Windows and Linux, with generation handled by your own NVIDIA GPU. It suits people who want to make AI video locally and see the clip develop as it renders.
6.1KUpdated 1 year agoApache-2.0
Web#Batch processing#Hugging Face integration#Multimodal input
LatentSync is an open-source AI lip-sync tool that edits a video's mouth movements to match supplied audio. It runs on your own GPU and suits video creators working with talking faces or virtual avatars, as well as researchers who want to train their own lip-sync models. The code uses the Apache 2.0 license.
12.6KUpdated 3 months ago
Web#Multimodal input#Quantization
HunyuanVideo is an AI video generation model for creators and developers who want to generate footage on their own hardware. Tencent provides model weights and inference code for text-to-video and image-to-video generation, alongside a hosted web experience. Local inference runs on your GPUs; the web offering runs through Tencent's service.
11KUpdated 9 months agoApache-2.0
#Hugging Face integration#LoRA#Multimodal input
LTX-Video is an AI video generation model for creators building controlled animations and developers adding video tools to their own products. You can run it locally or on your own servers using publicly available weights. The LTX family also offers a managed cloud API; local deployments can run in isolated environments without a cloud dependency.