
UNO generates images from reference images, with one model handling both a single subject and several subjects together. It's a local AI image generation framework from ByteDance for researchers and creators who want control over which subjects appear in a generated image and how consistently it preserves their appearance.
Its focus is subject consistency. Multiple reference images guide the output, so the model can place distinct subjects in the same image. UNO addresses attribute confusion, where a model mixes characteristics between subjects as more references enter the scene. Its universal rotary position embedding, called UnoPE, is designed to reduce that problem.
The framework builds on FLUX.1-dev and diffusion transformers. It trains progressively on single-subject data and then on multi-subject pairs, using synthesized images to support that progression. The repository includes inference and training code, plus a Gradio demo for running the model on your own hardware. Model checkpoints download from Hugging Face.
GPU memory is a practical consideration. The documented FP8 mode with offloading has a peak memory use of 16 GB; the project gives an RTX 3090 as its inference example.
The code is open source under Apache 2.0. The models use CC BY-NC 4.0, which restricts commercial use, and models based on FLUX.1-dev also remain subject to that base model's licensing terms.
Claim this page and we'll verify you by hand. UNO gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find UNO?Promote it
Something wrong or outdated on this page?
6.7KUpdated 2 years agoApache-2.0
#ControlNet#Hugging Face integration#Image-to-image
IP-Adapter lets you guide Stable Diffusion with a reference image while still using text to describe the result you want. It's for artists and developers who want image references in their local AI workflows. The main code and standard adapter weights use Apache 2.0; the separate FaceID variants are restricted to research use.
8.7KUpdated 11 months agoMIT
#Distributed execution#Hugging Face integration
VAR is an open-source image generation research toolkit that builds images from coarse representations to finer detail. It's for researchers and developers who want to run pretrained models on their own hardware or study an autoregressive alternative to diffusion. The code uses the MIT license.
4.3KUpdated 10 months ago
Linux · Web#ControlNet#Hugging Face integration#Image-to-image
HunyuanDiT is a text-to-image diffusion transformer for people who want to generate images on their own hardware using Chinese or English prompts. Its focus on Chinese language and cultural references makes it relevant for scenes drawn from poetry, idioms, food and architecture. It also handles long descriptions with details about subjects, clothing, lighting and artistic style.
34.1KUpdated 3 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
4.6KUpdated 2 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
4.3KUpdated 10 months agoMIT
Web#Hugging Face integration#Image-to-image#LoRA
ControlNet lets you guide Stable Diffusion with visual references, so a generated image can follow a sketch, a person's pose or the geometry of an existing scene. It's for artists and developers who need more control over image structure than a text prompt alone provides. The Python implementation runs on your own hardware and includes Gradio interfaces for its pretrained models.
Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.
OmniGen is a local AI image generation model that handles text prompts, reference images, and image editing within one model. It's for creators who want to reuse subjects across images and developers building image tools on their own hardware. The code is open source under the MIT license.