Favicon of Open-Sora

Open-Sora

Open source AI video generator you can run on your own GPUs, with image and text inputs, model training tools, and an Apache 2.0 license.

Screenshot of Open-Sora website

Open-Sora is an open source AI video generation project for developers, researchers, and creators who want to run and adapt a model on their own hardware. Its model focuses on turning reference images into video, with text prompts guiding the result. It also generates video directly from text. The code and Open-Sora 2.0 weights use Apache 2.0.

The same model handles both input types. For text prompts, a separate pipeline uses Flux to create an image before Open-Sora animates it. The default image stage uses FLUX.1-dev weights with their own license; Open-Sora's Apache license does not cover that component. The gallery includes realistic scenes and stylized animation, with examples of moving subjects and camera motion.

Users can influence movement through a motion score, and the project includes an evaluator for measuring motion in generated clips. Seed controls support reproducible results, while multiple samples let users compare different outputs from one prompt. Optional prompt refinement uses ChatGPT, so that feature sends prompts to an external service.

Open-Sora also provides tools and technical reports for training or fine-tuning a video model and training and evaluating its video autoencoder. Its Python code uses ColossalAI for work across multiple GPUs and supports offloading to reduce GPU memory pressure. Published performance tests use NVIDIA H100 and H800 GPUs, so the documented hardware context is server-class GPU computing.

Similar to Open-Sora