ControlNet lets you guide Stable Diffusion with visual references, so a generated image can follow a sketch, a person's pose or the geometry of an existing scene. It's for artists and developers who need more control over image structure than a text prompt alone provides. The Python implementation runs on your own hardware and includes Gradio interfaces for its pretrained models.
Different controls suit different tasks:
The supplied models work with Stable Diffusion 1.5, and the repository also supports Stable Diffusion 2. Guess Mode can generate images from a control map without a text prompt. Model weights and image detectors are available through Hugging Face, and a low-VRAM mode accommodates GPUs with 8 GB of memory.
ControlNet is open source under Apache 2.0. Developers can train it on their own paired images while keeping the original diffusion model's weights fixed. This approach learns the added visual control without overwriting the base model, and the project includes data annotation tools and training examples.
Claim this page and we'll verify you by hand. ControlNet gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find ControlNet?Promote it
Something wrong or outdated on this page?
4.3KUpdated 10 months ago
Linux · Web#ControlNet#Hugging Face integration#Image-to-image
HunyuanDiT is a text-to-image diffusion transformer for people who want to generate images on their own hardware using Chinese or English prompts. Its focus on Chinese language and cultural references makes it relevant for scenes drawn from poetry, idioms, food and architecture. It also handles long descriptions with details about subjects, clothing, lighting and artistic style.
4.6KUpdated 2 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
4.3KUpdated 10 months agoMIT
Web#Hugging Face integration#Image-to-image#LoRA
1.7KUpdated 2 years ago
Web#Image-to-image#Inpainting#Multimodal input
BrushNet adds text-guided image inpainting to pretrained diffusion models, letting you fill selected areas of an image while retaining the surrounding content. It's for developers and researchers who want to run image editing on their own hardware and work with an existing model's visual style.
1.4KUpdated 1 year agoApache-2.0
Web#Hugging Face integration#Image-to-image#Quantization
1KUpdated 4 months agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#LoRA
Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.
OmniGen is a local AI image generation model that handles text prompts, reference images, and image editing within one model. It's for creators who want to reuse subjects across images and developers building image tools on their own hardware. The code is open source under the MIT license.
UNO generates images from reference images, with one model handling both a single subject and several subjects together. It's a local AI image generation framework from ByteDance for researchers and creators who want control over which subjects appear in a generated image and how consistently it preserves their appearance.
Lumina-Image 2.0 is a local AI image generation framework for developers, researchers and people who want to generate images from text on their own hardware. It provides downloadable checkpoints, generation code and tools for adapting the model to your own image collections. The code uses the Apache 2.0 license.