CogView4 is a text-to-image model you can run on your own hardware, with support for Chinese and English prompts and Chinese text within generated images. It's aimed at developers and image creators who want local AI generation with native Chinese language support. The CogView4-6B model weights and repository code use Apache 2.0.
The model works with Hugging Face Diffusers, and a community wrapper brings it into ComfyUI workflows. A Gradio demo provides a browser interface. CogView4 uses GLM-4-9B to interpret prompts and accepts longer descriptions than CogView3-Plus, which supports English prompts. It supports square and rectangular images within its resolution limits.
Hardware demands are substantial. The provided inference example uses a CUDA GPU, and the project recommends at least 32GB of system RAM. Its batch-of-four tests report roughly 33–39GB of GPU memory without CPU offloading, or 13–14GB with offloading and a quantized text encoder. Those figures describe the tested workload, rather than a minimum for every image. BNB and TorchAO quantization support can reduce memory use.
Local inference runs on your hardware; Hugging Face Spaces, ModelScope Spaces and ZhipuAI MaaS provide hosted alternatives. The optional prompt-rewriting example sends prompts to ZhipuAI's API using GLM-4-Plus. For custom training, CogKit and finetrainers support LoRA and full fine-tuning outside this repository, with finetrainers supporting training on a single RTX 4090.
Claim this page and we'll verify you by hand. CogView4 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find CogView4?Promote it
Something wrong or outdated on this page?
4.6KUpdated 2 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.
1KUpdated 4 months agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#LoRA
17.8KUpdated 2 years agoMIT
Web#Batch processing#Hugging Face integration#Multimodal input
4.3KUpdated 10 months agoMIT
Web#Hugging Face integration#Image-to-image#LoRA
2.5KUpdated 1 year agoMIT
Web#Hugging Face integration
HiDream-I1 is an open-source text-to-image model for people who want to generate images on their own hardware or build image generation into a Python application. It uses MIT licensing and supports local inference through CUDA, making it an option for developers and creators with NVIDIA GPU hardware.
1.9KUpdated 2 years agoApache-2.0
Web#Hugging Face integration
PixArt-Sigma is a text-to-image diffusion model that supports generation at 2K and 4K resolutions on your own machine or server. It's aimed at developers and researchers who want pretrained models they can run themselves, along with code for training and adapting them. The Python code is open source under Apache 2.0.
Lumina-Image 2.0 is a local AI image generation framework for developers, researchers and people who want to generate images from text on their own hardware. It provides downloadable checkpoints, generation code and tools for adapting the model to your own image collections. The code uses the Apache 2.0 license.
Janus-Pro is a multimodal AI model from DeepSeek that answers questions about images and creates pictures from text prompts. It runs on your own hardware and suits developers and researchers who want both capabilities in one model. A local Gradio demo provides a browser interface, while Hugging Face hosts a separate online demo.
OmniGen is a local AI image generation model that handles text prompts, reference images, and image editing within one model. It's for creators who want to reuse subjects across images and developers building image tools on their own hardware. The code is open source under the MIT license.