
Stable Video Diffusion is an image-to-video model from Stability AI that animates a supplied still image into a short clip. It is for developers and researchers working with video generation on their own hardware. The SVD model generates 14 frames; SVD-XT is fine-tuned for 25 frames at 576×1024 resolution.
The official generative-models repository supplies Python inference scripts, a Streamlit video sampling interface and a community Gradio demo. Downloaded weights let you generate videos locally. The shared repository also contains other Stability models, but SDXL image generation and newer 3D models are separate from this entry.
The model has practical limits: outputs are short, motion can be limited, generated people and faces can be inaccurate, and text may not render legibly. It conditions on an image rather than accepting text prompts as direct instructions. GPU memory requirements depend on the runtime and generation settings.
The repository code uses MIT. The SVD-XT weights carry the Stability AI Community License with conditions for research, noncommercial and commercial use; the code license does not replace those model terms. Check the current license for your intended deployment.
Claim this page with an email at stability.ai. Stable Video Diffusion gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Stable Video Diffusion?Promote it
Something wrong or outdated on this page?
12.3KUpdated 2 years agoApache-2.0
Web#Hugging Face integration#LoRA#Multimodal input
AnimateDiff adds text-driven animation to personalized Stable Diffusion models without requiring separate training for each model. It's for artists and developers who want to generate motion while keeping the visual style of a chosen image model. The Python implementation runs locally and includes a Gradio browser interface.
13KUpdated 11 months agoApache-2.0
Windows · Web#Hugging Face integration#LoRA#Multimodal input
4.3KUpdated 6 months agoApache-2.0
Linux · Web#Hugging Face integration#Multimodal input
17.3KUpdated 11 months agoApache-2.0
Windows · Linux · Web#Hugging Face integration#Multimodal input
6.1KUpdated 1 year agoApache-2.0
Web#Batch processing#Hugging Face integration#Multimodal input
3.7KUpdated 11 months agoApache-2.0
Web#Hugging Face integration#LoRA
Mochi 1 is a text-to-video model for creators and developers who want to generate videos on their own hardware or adapt a model to their own footage. Genmo releases it under Apache 2.0, with downloadable weights and code for local use. Genmo also offers a hosted playground for trying the model in a browser.
CogVideoX is a family of downloadable video generation models for developers, researchers and creators who want to generate clips on their own hardware. It turns English text prompts into video, animates a supplied image and can continue an existing video. A local Gradio web interface provides a browser front end for generation.
EchoMimic turns a portrait image and an audio recording into an animated talking-head video on your own hardware. Its distinguishing feature is editable facial landmark control: you can drive animation with audio, landmarks, or audio combined with selected landmarks. It's intended for academic research and suits people comparing methods for speech animation and facial motion control.
FramePack is an open source desktop app for making videos from a still image and a written motion prompt. It runs on Windows and Linux, with generation handled by your own NVIDIA GPU. It suits people who want to make AI video locally and see the clip develop as it renders.
LatentSync is an open-source AI lip-sync tool that edits a video's mouth movements to match supplied audio. It runs on your own GPU and suits video creators working with talking faces or virtual avatars, as well as researchers who want to train their own lip-sync models. The code uses the Apache 2.0 license.