Favicon of Stable Video Diffusion

Stable Video Diffusion

An image-to-video model for animating still pictures locally, with Python and browser demos. Model weights use Stability AI Community terms.

Screenshot of Stable Video Diffusion website

Stable Video Diffusion is an image-to-video model from Stability AI that animates a supplied still image into a short clip. It is for developers and researchers working with video generation on their own hardware. The SVD model generates 14 frames; SVD-XT is fine-tuned for 25 frames at 576×1024 resolution.

The official generative-models repository supplies Python inference scripts, a Streamlit video sampling interface and a community Gradio demo. Downloaded weights let you generate videos locally. The shared repository also contains other Stability models, but SDXL image generation and newer 3D models are separate from this entry.

The model has practical limits: outputs are short, motion can be limited, generated people and faces can be inaccurate, and text may not render legibly. It conditions on an image rather than accepting text prompts as direct instructions. GPU memory requirements depend on the runtime and generation settings.

The repository code uses MIT. The SVD-XT weights carry the Stability AI Community License with conditions for research, noncommercial and commercial use; the code license does not replace those model terms. Check the current license for your intended deployment.

Similar to Stable Video Diffusion