
Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.
The model uses latent diffusion and supports detailed prompts with complex relationships between subjects. Its examples cover portraits, Chinese visual elements, and images containing lettering. Downloadable model weights and inference code let developers use it locally; the Python and PyTorch stack recommends CUDA for GPU acceleration. The code uses Apache 2.0, while the weights have separate model terms. Commercial use requires registration with the licensor, with additional conditions for very large services.
Beyond generating an image from a prompt, Kolors supports several ways to guide or edit the result:
Kolors works with ComfyUI, Gradio, and Hugging Face Diffusers, so it fits both visual image-generation workflows and Python applications. ComfyUI support makes it relevant to people who already build workflows around image guidance and editing, while Diffusers gives developers a supported route for incorporating the model into their own software.
Claim this page and we'll verify you by hand. Kolors gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Kolors?Promote it
Something wrong or outdated on this page?
4.3KUpdated 10 months agoMIT
Web#Hugging Face integration#Image-to-image#LoRA
OmniGen is a local AI image generation model that handles text prompts, reference images, and image editing within one model. It's for creators who want to reuse subjects across images and developers building image tools on their own hardware. The code is open source under the MIT license.
165.2KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Web#Batch processing#Code execution#Image-to-image
1KUpdated 4 months agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#LoRA
8.4KUpdated 8 months agoApache-2.0
Web#Image-to-image#LoRA#Multimodal input
1.9KUpdated 2 years agoApache-2.0
Web#Hugging Face integration
PixArt-Sigma is a text-to-image diffusion model that supports generation at 2K and 4K resolutions on your own machine or server. It's aimed at developers and researchers who want pretrained models they can run themselves, along with code for training and adapting them. The Python code is open source under Apache 2.0.
34.6KUpdated 1 day agoApache-2.0
macOS#ControlNet#Hugging Face integration#Image-to-image
Stable Diffusion web UI (AUTOMATIC1111) is a browser interface for generating and editing images with models running on your own hardware. It's for artists and anyone who wants control over prompts, models and image variations. The software is open source under AGPL-3.0.
Lumina-Image 2.0 is a local AI image generation framework for developers, researchers and people who want to generate images from text on their own hardware. It provides downloadable checkpoints, generation code and tools for adapting the model to your own image collections. The code uses the Apache 2.0 license.
Qwen-Image is an open-source image generation and editing model you can deploy locally. It's for developers and creators who want to generate images from text or revise existing pictures on their own hardware. Its text rendering capabilities, especially for Chinese, make it relevant for images that need readable lettering alongside visual content.
Diffusers is an open-source Python library for developers and researchers who want to run diffusion models on their own hardware or build generation features into an application. It uses PyTorch and supports image, video and audio generation. The library is licensed under Apache 2.0 and supports Apple Silicon.