
Pyramid Flow is an open-source AI video generator for people who want to create clips on their own hardware, and for researchers working on video models. It turns text prompts into video or animates a supplied image with guidance from text. The Python code uses PyTorch and carries MIT. Model licenses are separate: the SD3-derived weights use the Stability AI Community License. Check the chosen checkpoint before deployment.
The model can produce clips up to 10 seconds long at 768p and 24 frames per second. Text-guided image animation supports five-second clips. You can choose between miniFLUX and SD3 model variants; miniFLUX focuses on better human anatomy and more stable motion. A local Gradio interface lets you work in a browser, while Hugging Face Spaces and Google Colab provide cloud alternatives.
Hardware needs matter here. CPU offloading reduces GPU memory use to below 12GB, or below 8GB with sequential offloading, at the cost of speed. Apple Silicon can run the model through the MPS backend. Multi-GPU inference for the SD3 variant spreads the workload across two or four GPUs to reduce memory use per GPU and speed up generation.
Its distinguishing approach is pyramidal flow matching: generation works across different resolutions and noise levels rather than doing all the work at full resolution. The model generates video progressively and was trained using only open-source datasets. The repository also includes training and fine-tuning code, with at least eight A100 GPUs required for VAE training or DiT fine-tuning.
Claim this page and we'll verify you by hand. Pyramid Flow gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Pyramid Flow?Promote it
Something wrong or outdated on this page?
29.9KUpdated 6 months agoApache-2.0
#Hugging Face integration#Multimodal input
Open-Sora is an open source AI video generation project for developers, researchers, and creators who want to run and adapt a model on their own hardware. Its model focuses on turning reference images into video, with text prompts guiding the result. It also generates video directly from text. The code and Open-Sora 2.0 weights use Apache 2.0.
12.3KUpdated 2 years agoApache-2.0
Web#Hugging Face integration#LoRA#Multimodal input
13KUpdated 11 months agoApache-2.0
Windows · Web#Hugging Face integration#LoRA#Multimodal input
17.3KUpdated 11 months agoApache-2.0
Windows · Linux · Web#Hugging Face integration#Multimodal input
6.1KUpdated 1 year agoApache-2.0
Web#Batch processing#Hugging Face integration#Multimodal input
7.6KUpdated 8 months ago
Web#Hugging Face integration#Multimodal input
AnimateDiff adds text-driven animation to personalized Stable Diffusion models without requiring separate training for each model. It's for artists and developers who want to generate motion while keeping the visual style of a chosen image model. The Python implementation runs locally and includes a Gradio browser interface.
CogVideoX is a family of downloadable video generation models for developers, researchers and creators who want to generate clips on their own hardware. It turns English text prompts into video, animates a supplied image and can continue an existing video. A local Gradio web interface provides a browser front end for generation.
FramePack is an open source desktop app for making videos from a still image and a written motion prompt. It runs on Windows and Linux, with generation handled by your own NVIDIA GPU. It suits people who want to make AI video locally and see the clip develop as it renders.
LatentSync is an open-source AI lip-sync tool that edits a video's mouth movements to match supplied audio. It runs on your own GPU and suits video creators working with talking faces or virtual avatars, as well as researchers who want to train their own lip-sync models. The code uses the Apache 2.0 license.
SkyReels-V2 is a family of downloadable video models for developers and researchers who want to run generation on their own GPUs. It supports text-to-video, image-to-video and extended video generation. Its autoregressive Diffusion Forcing architecture generates longer sequences by continuing video segments.