
Sana is an open-source framework for running image and video generation on your own hardware, with image models small enough for laptop GPUs. It's aimed at creators who want local AI generation and developers who need training and inference pipelines for their own models. The code uses the Apache 2.0 license.
Its text-to-image models produce images up to 4K resolution. The Sana-0.6B model can run on a laptop GPU with 16GB of memory, while 4-bit quantization brings image generation below 8GB of VRAM. Those laptop claims apply to image generation, rather than every video model in the family. Hosted demos and a Replicate API provide separate ways to try generation without running the models locally.
Sana works with ComfyUI and Hugging Face Diffusers. LoRA and DreamBooth support let users customize image models, and ControlNet adds control over generated images beyond the text prompt. SANA-Sprint offers image generation in one or a few steps.
The efficiency comes partly from an autoencoder that compresses images more deeply and linear attention that reduces the work needed at high resolutions. Gemma handles text encoding to help the image models follow prompts, with its own Gemma terms applying to that component.
The video tools cover video generation, controllable world generation, and streaming video editing. SANA-WM supports camera movement control in generated worlds, while SANA-Streaming edits video as it streams and maintains consistency across frames.
Claim this page and we'll verify you by hand. Sana gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Sana?Promote it
Something wrong or outdated on this page?
13KUpdated 11 months agoApache-2.0
Windows · Web#Hugging Face integration#LoRA#Multimodal input
CogVideoX is a family of downloadable video generation models for developers, researchers and creators who want to generate clips on their own hardware. It turns English text prompts into video, animates a supplied image and can continue an existing video. A local Gradio web interface provides a browser front end for generation.
13.2KUpdated 2 days agoApache-2.0
#ControlNet#Image-to-image#Inpainting
11KUpdated 9 months agoApache-2.0
#Hugging Face integration#LoRA#Multimodal input
12.3KUpdated 2 years agoApache-2.0
Web#Hugging Face integration#LoRA#Multimodal input
13.6KUpdated 3 years agoAGPL-3.0
macOS#ControlNet#Image-to-image#Inpainting
12.6KUpdated 3 months ago
Web#Multimodal input#Quantization
HunyuanVideo is an AI video generation model for creators and developers who want to generate footage on their own hardware. Tencent provides model weights and inference code for text-to-video and image-to-video generation, alongside a hosted web experience. Local inference runs on your GPUs; the web offering runs through Tencent's service.
DiffSynth-Studio is a Python diffusion model engine for developers and researchers who want to generate media and train models on their own hardware. It supports large models on consumer GPUs through memory offloading and quantization, with inference and training in the same framework. It's open source under Apache 2.0.
LTX-Video is an AI video generation model for creators building controlled animations and developers adding video tools to their own products. You can run it locally or on your own servers using publicly available weights. The LTX family also offers a managed cloud API; local deployments can run in isolated environments without a cloud dependency.
AnimateDiff adds text-driven animation to personalized Stable Diffusion models without requiring separate training for each model. It's for artists and developers who want to generate motion while keeping the visual style of a chosen image model. The Python implementation runs locally and includes a Gradio browser interface.
DiffusionBee is an open-source AI art app for Mac users who want to generate and edit images on their own computer. It runs Stable Diffusion offline, with image generation processed on the device. Model downloads require network access, and optional image uploads can send images externally. Its visual interface suits artists and designers who want local image tools without working through code.