Favicon of CogView4

CogView4

A local text-to-image model for Chinese and English prompts, with Diffusers support, a community ComfyUI wrapper, and Apache 2.0 code.

CogView4 is a text-to-image model you can run on your own hardware, with support for Chinese and English prompts and Chinese text within generated images. It's aimed at developers and image creators who want local AI generation with native Chinese language support. The CogView4-6B model weights and repository code use Apache 2.0.

The model works with Hugging Face Diffusers, and a community wrapper brings it into ComfyUI workflows. A Gradio demo provides a browser interface. CogView4 uses GLM-4-9B to interpret prompts and accepts longer descriptions than CogView3-Plus, which supports English prompts. It supports square and rectangular images within its resolution limits.

Hardware demands are substantial. The provided inference example uses a CUDA GPU, and the project recommends at least 32GB of system RAM. Its batch-of-four tests report roughly 33–39GB of GPU memory without CPU offloading, or 13–14GB with offloading and a quantized text encoder. Those figures describe the tested workload, rather than a minimum for every image. BNB and TorchAO quantization support can reduce memory use.

Local inference runs on your hardware; Hugging Face Spaces, ModelScope Spaces and ZhipuAI MaaS provide hosted alternatives. The optional prompt-rewriting example sends prompts to ZhipuAI's API using GLM-4-Plus. For custom training, CogKit and finetrainers support LoRA and full fine-tuning outside this repository, with finetrainers supporting training on a single RTX 4090.

Similar to CogView4