Favicon of HiDream-I1

HiDream-I1

An MIT-licensed image generation model you can run locally with CUDA, with full and distilled variants, Diffusers support, and a Gradio interface.

Screenshot of HiDream-I1 website

HiDream-I1 is an open-source text-to-image model for people who want to generate images on their own hardware or build image generation into a Python application. It uses MIT licensing and supports local inference through CUDA, making it an option for developers and creators with NVIDIA GPU hardware.

The model comes in Full, Dev, and Fast variants. Full uses more generation steps, while the distilled Dev and Fast models use fewer. That gives users a choice between the full model and alternatives designed to shorten generation runs, without moving inference to a hosted service.

HiDream-I1 works with Hugging Face Diffusers, so it can fit into an existing Python image-generation workflow. It also includes a Gradio demo for generating images through an interactive interface. The supplied inference code downloads meta-llama/Llama-3.1-8B-Instruct model files, so acquiring those files requires a Hugging Face account that accepts the Llama 3.1 terms. Its text encoders and VAE retain their own licenses, separate from the MIT-licensed HiDream model.

Its evaluations cover prompt following and visual preference across animation, concept art, painting, and photography. In the reported DPG-Bench results, it scores above Flux.1-dev, DALL-E 3, and SD3-Medium overall; that benchmark tests how images reflect objects, attributes, and relationships in a prompt. The project also reports higher average human-preference scores than those models in its HPS evaluation.

Similar to HiDream-I1