Favicon of Stable Diffusion 1.5

Stable Diffusion 1.5

A local AI image generation model with text-guided image editing, Diffusers support, and a reference implementation requiring at least 10GB of GPU VRAM.

Screenshot of Stable Diffusion 1.5 website

Stable Diffusion 1.5 is an AI image generation model for creators and developers who want to generate images on their own hardware. It turns text prompts into images and supports text-guided changes to existing pictures, including turning rough sketches into detailed artwork. Local inference keeps that image generation work on your machine.

The model supports image-to-image generation and upscaling as well as generation from text alone. When modifying an image, you can control how closely the result follows the original: stronger changes allow more variation but can lose the original image's content. That makes it useful for exploring visual ideas from an existing drawing or picture.

Stable Diffusion uses latent diffusion, which generates images in a compressed representation. The supplied implementation runs on a GPU with at least 10GB of VRAM, and the model was trained on 512×512 images. Developers can use it through the Hugging Face Diffusers library or the project's reference code.

The downloadable weights use the CreativeML OpenRAIL-M license, which permits commercial use subject to its use restrictions. The reference generation code includes a safety checker intended to reduce explicit outputs and an invisible watermark to identify generated images. Its training data comes from subsets of LAION-5B, and the model can reproduce biases present in that data.

Similar to Stable Diffusion 1.5