
Stable Diffusion 1.5 is an AI image generation model for creators and developers who want to generate images on their own hardware. It turns text prompts into images and supports text-guided changes to existing pictures, including turning rough sketches into detailed artwork. Local inference keeps that image generation work on your machine.
The model supports image-to-image generation and upscaling as well as generation from text alone. When modifying an image, you can control how closely the result follows the original: stronger changes allow more variation but can lose the original image's content. That makes it useful for exploring visual ideas from an existing drawing or picture.
Stable Diffusion uses latent diffusion, which generates images in a compressed representation. The supplied implementation runs on a GPU with at least 10GB of VRAM, and the model was trained on 512×512 images. Developers can use it through the Hugging Face Diffusers library or the project's reference code.
The downloadable weights use the CreativeML OpenRAIL-M license, which permits commercial use subject to its use restrictions. The reference generation code includes a safety checker intended to reduce explicit outputs and an invisible watermark to identify generated images. Its training data comes from subsets of LAION-5B, and the model can reproduce biases present in that data.
Claim this page with an email at stability.ai. Stable Diffusion 1.5 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Stable Diffusion 1.5?Promote it
Something wrong or outdated on this page?
4.6KUpdated 2 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.
4.3KUpdated 10 months agoMIT
Web#Hugging Face integration#Image-to-image#LoRA
26KUpdated 1 year agoApache-2.0
#Image-to-image#Inpainting#LoRA
8.4KUpdated 8 months agoApache-2.0
Web#Image-to-image#LoRA#Multimodal input
34.6KUpdated 1 day agoApache-2.0
macOS#ControlNet#Hugging Face integration#Image-to-image
1.5KUpdated 2 years agoMIT
#ControlNet#Hugging Face integration#Image-to-image
OmniGen is a local AI image generation model that handles text prompts, reference images, and image editing within one model. It's for creators who want to reuse subjects across images and developers building image tools on their own hardware. The code is open source under the MIT license.
FLUX.1 is a family of image models for people who want to generate or edit images on their own infrastructure. Black Forest Labs provides Python inference code for its open-weight models and a separate hosted API. Local inference runs on your hardware; API requests go to Black Forest Labs.
Qwen-Image is an open-source image generation and editing model you can deploy locally. It's for developers and creators who want to generate images from text or revise existing pictures on their own hardware. Its text rendering capabilities, especially for Chinese, make it relevant for images that need readable lettering alongside visual content.
Diffusers is an open-source Python library for developers and researchers who want to run diffusion models on their own hardware or build generation features into an application. It uses PyTorch and supports image, video and audio generation. The library is licensed under Apache 2.0 and supports Apple Silicon.
Stable Diffusion 3.5 is a family of text-to-image models for people building image tools or producing visual work on their own infrastructure. It generates photography, paintings, line art and 3D-style images from prompts, with an emphasis on following the requested subject and composition.