Favicon of Sana

Sana

An open-source image and video generation framework with 4K text-to-image models, laptop GPU support, and ComfyUI and Diffusers integrations.

Screenshot of Sana website

Sana is an open-source framework for running image and video generation on your own hardware, with image models small enough for laptop GPUs. It's aimed at creators who want local AI generation and developers who need training and inference pipelines for their own models. The code uses the Apache 2.0 license.

Its text-to-image models produce images up to 4K resolution. The Sana-0.6B model can run on a laptop GPU with 16GB of memory, while 4-bit quantization brings image generation below 8GB of VRAM. Those laptop claims apply to image generation, rather than every video model in the family. Hosted demos and a Replicate API provide separate ways to try generation without running the models locally.

Sana works with ComfyUI and Hugging Face Diffusers. LoRA and DreamBooth support let users customize image models, and ControlNet adds control over generated images beyond the text prompt. SANA-Sprint offers image generation in one or a few steps.

The efficiency comes partly from an autoencoder that compresses images more deeply and linear attention that reduces the work needed at high resolutions. Gemma handles text encoding to help the image models follow prompts, with its own Gemma terms applying to that component.

The video tools cover video generation, controllable world generation, and streaming video editing. SANA-WM supports camera movement control in generated worlds, while SANA-Streaming edits video as it streams and maintains consistency across frames.

Similar to Sana