
HunyuanDiT is a text-to-image diffusion transformer for people who want to generate images on their own hardware using Chinese or English prompts. Its focus on Chinese language and cultural references makes it relevant for scenes drawn from poetry, idioms, food and architecture. It also handles long descriptions with details about subjects, clothing, lighting and artistic style.
You can refine images conversationally. The companion DialogGen model interprets follow-up requests and turns the conversation into a revised image prompt, so you can ask for changes while retaining context. Tencent also offers image generation through its hosted Hunyuan Bot; the downloadable models and inference code support local use.
The model works with ComfyUI and Hugging Face Diffusers, and includes Gradio demos. For more control over the result, it supports ControlNet guidance using pose, edges or depth, plus IP-adapter for image references. Training code and LoRA support let developers adapt the model to their own material. Distilled and TensorRT variants provide other inference options.
Local inference requires an NVIDIA GPU with CUDA support, and Linux is the tested operating system. An optimized Diffusers implementation can run below 6GB of GPU memory on supported NVIDIA Ampere cards and later examples such as the RTX 4080 and 4090. The standard model uses more memory, and running DialogGen alongside it raises the requirement further.
Claim this page with an email at dit.hunyuan.tencent.com. HunyuanDiT gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find HunyuanDiT?Promote it
Something wrong or outdated on this page?
1KUpdated 4 months agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#LoRA
Lumina-Image 2.0 is a local AI image generation framework for developers, researchers and people who want to generate images from text on their own hardware. It provides downloadable checkpoints, generation code and tools for adapting the model to your own image collections. The code uses the Apache 2.0 license.
34.1KUpdated 3 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
4.6KUpdated 2 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
4.3KUpdated 10 months agoMIT
Web#Hugging Face integration#Image-to-image#LoRA
2.3KUpdated 5 months agoApache-2.0
Linux#Distributed execution#Image-to-image#LoRA
Step1X Edit is an open-source image editing model for developers and researchers who want to edit existing pictures with written instructions on their own GPU hardware. It handles changes to individual objects as well as the overall appearance of an image. The code uses the Apache 2.0 license.
1.4KUpdated 1 year agoApache-2.0
Web#Hugging Face integration#Image-to-image#Quantization
ControlNet lets you guide Stable Diffusion with visual references, so a generated image can follow a sketch, a person's pose or the geometry of an existing scene. It's for artists and developers who need more control over image structure than a text prompt alone provides. The Python implementation runs on your own hardware and includes Gradio interfaces for its pretrained models.
Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.
OmniGen is a local AI image generation model that handles text prompts, reference images, and image editing within one model. It's for creators who want to reuse subjects across images and developers building image tools on their own hardware. The code is open source under the MIT license.
UNO generates images from reference images, with one model handling both a single subject and several subjects together. It's a local AI image generation framework from ByteDance for researchers and creators who want control over which subjects appear in a generated image and how consistently it preserves their appearance.