VAR is an open-source image generation research toolkit that builds images from coarse representations to finer detail. It's for researchers and developers who want to run pretrained models on their own hardware or study an autoregressive alternative to diffusion. The code uses the MIT license.
Its defining approach is next-scale prediction: the model predicts a finer image representation at each stage, rather than generating image tokens one at a time in raster order. It uses GPT-style transformers, with downloadable VAR checkpoints and an accompanying image autoencoder. The implementation runs in Python with PyTorch, and Jupyter notebooks demonstrate image sampling and zero-shot editing.
The supplied models generate images at 256×256 or 512×512 resolution. Different model sizes let researchers compare generation quality and computational cost. Sampling controls adjust the balance between visual quality and diversity, and the evaluation workflow supports FID, Inception Score, precision and recall.
VAR also includes ImageNet training code, distributed training support and automatic recovery from saved checkpoints after an interruption. Training checkpoints and logs stay in a local output folder. Optional FlashAttention and xformers support can speed up attention computation.
Alongside the downloadable code and weights, a hosted browser demo provides interactive text-to-image generation. Infinity is a separate related project for text-to-image research; VAR's repository supplies the image generation models and training implementation described here.
Claim this page and we'll verify you by hand. VAR gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find VAR?Promote it
Something wrong or outdated on this page?
6.7KUpdated 2 years agoApache-2.0
#ControlNet#Hugging Face integration#Image-to-image
IP-Adapter lets you guide Stable Diffusion with a reference image while still using text to describe the result you want. It's for artists and developers who want image references in their local AI workflows. The main code and standard adapter weights use Apache 2.0; the separate FaceID variants are restricted to research use.
1.4KUpdated 1 year agoApache-2.0
Web#Hugging Face integration#Image-to-image#Quantization
34.1KUpdated 3 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
34.6KUpdated 18 hours agoApache-2.0
macOS#ControlNet#Hugging Face integration#Image-to-image
4.3KUpdated 10 months ago
Linux · Web#ControlNet#Hugging Face integration#Image-to-image
HunyuanDiT is a text-to-image diffusion transformer for people who want to generate images on their own hardware using Chinese or English prompts. Its focus on Chinese language and cultural references makes it relevant for scenes drawn from poetry, idioms, food and architecture. It also handles long descriptions with details about subjects, clothing, lighting and artistic style.
4.6KUpdated 2 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
UNO generates images from reference images, with one model handling both a single subject and several subjects together. It's a local AI image generation framework from ByteDance for researchers and creators who want control over which subjects appear in a generated image and how consistently it preserves their appearance.
ControlNet lets you guide Stable Diffusion with visual references, so a generated image can follow a sketch, a person's pose or the geometry of an existing scene. It's for artists and developers who need more control over image structure than a text prompt alone provides. The Python implementation runs on your own hardware and includes Gradio interfaces for its pretrained models.
Diffusers is an open-source Python library for developers and researchers who want to run diffusion models on their own hardware or build generation features into an application. It uses PyTorch and supports image, video and audio generation. The library is licensed under Apache 2.0 and supports Apple Silicon.
Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.