
PixArt-Sigma is a text-to-image diffusion model that supports generation at 2K and 4K resolutions on your own machine or server. It's aimed at developers and researchers who want pretrained models they can run themselves, along with code for training and adapting them. The Python code is open source under Apache 2.0.
The model builds on PixArt-α's diffusion transformer architecture. It uses a T5 text encoder and the SDXL variational autoencoder, with support for longer prompts than PixArt-α. Its weak-to-strong training approach is the focus of the accompanying research.
For image generation, you can use the PyTorch inference code or Hugging Face Diffusers through its PixArtSigmaPipeline. A Gradio interface provides a browser front end for local use, and the demo can also run on a self-hosted server. Model files are available to download, so generation can run on hardware you control rather than through a hosted image service.
The repository includes training and sampling code as well as pretrained checkpoints at several resolutions, including 1024 pixels and 2K. Training doesn't require a separate feature-extraction stage. It also supports converting PyTorch .pth checkpoints into the Diffusers format.
The pretrained model weights use CreativeML Open RAIL++-M terms, separate from the Apache-licensed Python code.
Claim this page and we'll verify you by hand. PixArt-Sigma gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find PixArt-Sigma?Promote it
Something wrong or outdated on this page?
4.6KUpdated 2 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.
1KUpdated 4 months agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#LoRA
4.3KUpdated 10 months agoMIT
Web#Hugging Face integration#Image-to-image#LoRA
9.2KUpdated 2 weeks agoApache-2.0
#ControlNet#LoRA#Multimodal input
12.2KUpdated 3 days agoMIT
macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#LoRA
1.1KUpdated 2 years agoApache-2.0
Web#Batch processing#Hugging Face integration#LoRA
Lumina-Image 2.0 is a local AI image generation framework for developers, researchers and people who want to generate images from text on their own hardware. It provides downloadable checkpoints, generation code and tools for adapting the model to your own image collections. The code uses the Apache 2.0 license.
OmniGen is a local AI image generation model that handles text prompts, reference images, and image editing within one model. It's for creators who want to reuse subjects across images and developers building image tools on their own hardware. The code is open source under the MIT license.
Sana is an open-source framework for running image and video generation on your own hardware, with image models small enough for laptop GPUs. It's aimed at creators who want local AI generation and developers who need training and inference pipelines for their own models. The code uses the Apache 2.0 license.
AI Toolkit (ostris) is an MIT-licensed training suite for people who want to fine-tune image and video models on their own hardware or a self-hosted server. It targets consumer NVIDIA GPUs and runs on Linux and Windows, including ARM64 Linux systems such as DGX Spark. An experimental installer also supports Apple Silicon Macs. GPU memory needs depend on the model and training task.
CogView4 is a text-to-image model you can run on your own hardware, with support for Chinese and English prompts and Chinese text within generated images. It's aimed at developers and image creators who want local AI generation with native Chinese language support. The CogView4-6B model weights and repository code use Apache 2.0.