
HiDream-I1 is an open-source text-to-image model for people who want to generate images on their own hardware or build image generation into a Python application. It uses MIT licensing and supports local inference through CUDA, making it an option for developers and creators with NVIDIA GPU hardware.
The model comes in Full, Dev, and Fast variants. Full uses more generation steps, while the distilled Dev and Fast models use fewer. That gives users a choice between the full model and alternatives designed to shorten generation runs, without moving inference to a hosted service.
HiDream-I1 works with Hugging Face Diffusers, so it can fit into an existing Python image-generation workflow. It also includes a Gradio demo for generating images through an interactive interface. The supplied inference code downloads meta-llama/Llama-3.1-8B-Instruct model files, so acquiring those files requires a Hugging Face account that accepts the Llama 3.1 terms. Its text encoders and VAE retain their own licenses, separate from the MIT-licensed HiDream model.
Its evaluations cover prompt following and visual preference across animation, concept art, painting, and photography. In the reported DPG-Bench results, it scores above Flux.1-dev, DALL-E 3, and SD3-Medium overall; that benchmark tests how images reflect objects, attributes, and relationships in a prompt. The project also reports higher average human-preference scores than those models in its HPS evaluation.
Claim this page with an email at hidream.ai. HiDream-I1 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find HiDream-I1?Promote it
Something wrong or outdated on this page?
1.1KUpdated 2 years agoApache-2.0
Web#Batch processing#Hugging Face integration#LoRA
CogView4 is a text-to-image model you can run on your own hardware, with support for Chinese and English prompts and Chinese text within generated images. It's aimed at developers and image creators who want local AI generation with native Chinese language support. The CogView4-6B model weights and repository code use Apache 2.0.
17.8KUpdated 2 years agoMIT
Web#Batch processing#Hugging Face integration#Multimodal input
4.6KUpdated 2 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
1KUpdated 4 months agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#LoRA
4.3KUpdated 10 months agoMIT
Web#Hugging Face integration#Image-to-image#LoRA
1.9KUpdated 2 years agoApache-2.0
Web#Hugging Face integration
PixArt-Sigma is a text-to-image diffusion model that supports generation at 2K and 4K resolutions on your own machine or server. It's aimed at developers and researchers who want pretrained models they can run themselves, along with code for training and adapting them. The Python code is open source under Apache 2.0.
Janus-Pro is a multimodal AI model from DeepSeek that answers questions about images and creates pictures from text prompts. It runs on your own hardware and suits developers and researchers who want both capabilities in one model. A local Gradio demo provides a browser interface, while Hugging Face hosts a separate online demo.
Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.
Lumina-Image 2.0 is a local AI image generation framework for developers, researchers and people who want to generate images from text on their own hardware. It provides downloadable checkpoints, generation code and tools for adapting the model to your own image collections. The code uses the Apache 2.0 license.
OmniGen is a local AI image generation model that handles text prompts, reference images, and image editing within one model. It's for creators who want to reuse subjects across images and developers building image tools on their own hardware. The code is open source under the MIT license.