
BrushNet adds text-guided image inpainting to pretrained diffusion models, letting you fill selected areas of an image while retaining the surrounding content. It's for developers and researchers who want to run image editing on their own hardware and work with an existing model's visual style.
It accepts an image, a mask marking the area to edit, and a text prompt. Its dual-branch design processes the masked image separately from the diffusion model's noisy image representation, giving generation detailed guidance from the original picture. An adjustable control strength changes how closely the result follows that guidance.
BrushNet provides checkpoints for Stable Diffusion 1.5 and Stable Diffusion XL, with separate training for object-shaped segmentation masks and more general, irregular masks. It also works with community models such as DreamShaper, epiCRealism, MeinaMix and Realistic Vision. Examples cover photographs, anime, pencil drawings and watercolor, so the choice of base model matters for the look of the edit.
The Python and PyTorch implementation includes inference, a Gradio demo, and training on custom datasets. BrushData supplies segmentation annotations for training, while BrushBench pairs images with human-annotated masks and captions for evaluation. The pretrained model targets general scenes; specialized uses such as product displays or virtual try-on may require training on your own data.
Claim this page and we'll verify you by hand. BrushNet gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find BrushNet?Promote it
Something wrong or outdated on this page?
4.6KUpdated 2 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.
4.3KUpdated 10 months agoMIT
Web#Hugging Face integration#Image-to-image#LoRA
2.3KUpdated 5 months agoApache-2.0
Linux#Distributed execution#Image-to-image#LoRA
Step1X Edit is an open-source image editing model for developers and researchers who want to edit existing pictures with written instructions on their own GPU hardware. It handles changes to individual objects as well as the overall appearance of an image. The code uses the Apache 2.0 license.
34.1KUpdated 3 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
4.3KUpdated 10 months ago
Linux · Web#ControlNet#Hugging Face integration#Image-to-image
HunyuanDiT is a text-to-image diffusion transformer for people who want to generate images on their own hardware using Chinese or English prompts. Its focus on Chinese language and cultural references makes it relevant for scenes drawn from poetry, idioms, food and architecture. It also handles long descriptions with details about subjects, clothing, lighting and artistic style.
8.4KUpdated 8 months agoApache-2.0
Web#Image-to-image#LoRA#Multimodal input
OmniGen is a local AI image generation model that handles text prompts, reference images, and image editing within one model. It's for creators who want to reuse subjects across images and developers building image tools on their own hardware. The code is open source under the MIT license.
ControlNet lets you guide Stable Diffusion with visual references, so a generated image can follow a sketch, a person's pose or the geometry of an existing scene. It's for artists and developers who need more control over image structure than a text prompt alone provides. The Python implementation runs on your own hardware and includes Gradio interfaces for its pretrained models.
Qwen-Image is an open-source image generation and editing model you can deploy locally. It's for developers and creators who want to generate images from text or revise existing pictures on their own hardware. Its text rendering capabilities, especially for Chinese, make it relevant for images that need readable lettering alongside visual content.