Favicon of Kolors

Kolors

A local image model with Chinese and English prompts, text rendering, ComfyUI support, Apache 2.0 code, and separate model-weight terms.

Screenshot of Kolors website

Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.

The model uses latent diffusion and supports detailed prompts with complex relationships between subjects. Its examples cover portraits, Chinese visual elements, and images containing lettering. Downloadable model weights and inference code let developers use it locally; the Python and PyTorch stack recommends CUDA for GPU acceleration. The code uses Apache 2.0, while the weights have separate model terms. Commercial use requires registration with the licensor, with additional conditions for very large services.

Beyond generating an image from a prompt, Kolors supports several ways to guide or edit the result:

  • IP-Adapter-Plus uses reference images to guide generation, while IP-Adapter-FaceID-Plus provides face-based guidance.
  • ControlNet accepts edge, depth, and pose guidance for control over composition and structure.
  • Inpainting edits selected image regions. Dreambooth-LoRA training and inference support model customization.

Kolors works with ComfyUI, Gradio, and Hugging Face Diffusers, so it fits both visual image-generation workflows and Python applications. ComfyUI support makes it relevant to people who already build workflows around image guidance and editing, while Diffusers gives developers a supported route for incorporating the model into their own software.

Similar to Kolors