Favicon of PixArt-Sigma

PixArt-Sigma

An open-source text-to-image model for local or self-hosted generation, with 4K output, PyTorch training code and Hugging Face Diffusers support.

Screenshot of PixArt-Sigma website

PixArt-Sigma is a text-to-image diffusion model that supports generation at 2K and 4K resolutions on your own machine or server. It's aimed at developers and researchers who want pretrained models they can run themselves, along with code for training and adapting them. The Python code is open source under Apache 2.0.

The model builds on PixArt-α's diffusion transformer architecture. It uses a T5 text encoder and the SDXL variational autoencoder, with support for longer prompts than PixArt-α. Its weak-to-strong training approach is the focus of the accompanying research.

For image generation, you can use the PyTorch inference code or Hugging Face Diffusers through its PixArtSigmaPipeline. A Gradio interface provides a browser front end for local use, and the demo can also run on a self-hosted server. Model files are available to download, so generation can run on hardware you control rather than through a hosted image service.

The repository includes training and sampling code as well as pretrained checkpoints at several resolutions, including 1024 pixels and 2K. Training doesn't require a separate feature-extraction stage. It also supports converting PyTorch .pth checkpoints into the Diffusers format.

The pretrained model weights use CreativeML Open RAIL++-M terms, separate from the Apache-licensed Python code.

Similar to PixArt-Sigma