
Stable Diffusion XL is an image generation model for creators and developers who want to generate images on their own hardware or servers. It supports text-to-image generation and image-to-image sampling, so you can start with a written prompt or an existing image.
The model family includes SDXL-base and SDXL-refiner. Stability AI's Python code provides a Streamlit demo for both generation modes, alongside training examples for developers who want to build or adapt diffusion models. The code uses PyTorch, with modular components that let developers change parts of the model and training setup.
Licensing has two separate parts. The repository's code is open source under the MIT License, while the SDXL weights use the CreativeML Open RAIL++-M license. The code license doesn't replace the terms attached to the model weights.
Self-hosting lets you run the models in your own environment and control your data. Stability AI also offers hosted applications and media generation APIs, which run as managed services rather than on your machine.
Images generated through the supplied code contain an invisible watermark. The repository includes a detection script for that watermark.
Claim this page with an email at stability.ai. Stable Diffusion XL gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Stable Diffusion XL?Promote it
Something wrong or outdated on this page?
4.6KUpdated 2 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.
4.3KUpdated 10 months agoMIT
Web#Hugging Face integration#Image-to-image#LoRA
1.1KUpdated 2 years agoApache-2.0
Web#Batch processing#Hugging Face integration#LoRA
2.5KUpdated 1 year agoMIT
Web#Hugging Face integration
HiDream-I1 is an open-source text-to-image model for people who want to generate images on their own hardware or build image generation into a Python application. It uses MIT licensing and supports local inference through CUDA, making it an option for developers and creators with NVIDIA GPU hardware.
17.8KUpdated 2 years agoMIT
Web#Batch processing#Hugging Face integration#Multimodal input
1KUpdated 4 months agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#LoRA
OmniGen is a local AI image generation model that handles text prompts, reference images, and image editing within one model. It's for creators who want to reuse subjects across images and developers building image tools on their own hardware. The code is open source under the MIT license.
CogView4 is a text-to-image model you can run on your own hardware, with support for Chinese and English prompts and Chinese text within generated images. It's aimed at developers and image creators who want local AI generation with native Chinese language support. The CogView4-6B model weights and repository code use Apache 2.0.
Janus-Pro is a multimodal AI model from DeepSeek that answers questions about images and creates pictures from text prompts. It runs on your own hardware and suits developers and researchers who want both capabilities in one model. A local Gradio demo provides a browser interface, while Hugging Face hosts a separate online demo.
Lumina-Image 2.0 is a local AI image generation framework for developers, researchers and people who want to generate images from text on their own hardware. It provides downloadable checkpoints, generation code and tools for adapting the model to your own image collections. The code uses the Apache 2.0 license.