Favicon of VAR

VAR

An open-source image generation research toolkit that runs downloaded models in PyTorch, with MIT-licensed code and ImageNet training support.

VAR is an open-source image generation research toolkit that builds images from coarse representations to finer detail. It's for researchers and developers who want to run pretrained models on their own hardware or study an autoregressive alternative to diffusion. The code uses the MIT license.

Its defining approach is next-scale prediction: the model predicts a finer image representation at each stage, rather than generating image tokens one at a time in raster order. It uses GPT-style transformers, with downloadable VAR checkpoints and an accompanying image autoencoder. The implementation runs in Python with PyTorch, and Jupyter notebooks demonstrate image sampling and zero-shot editing.

The supplied models generate images at 256×256 or 512×512 resolution. Different model sizes let researchers compare generation quality and computational cost. Sampling controls adjust the balance between visual quality and diversity, and the evaluation workflow supports FID, Inception Score, precision and recall.

VAR also includes ImageNet training code, distributed training support and automatic recovery from saved checkpoints after an interruption. Training checkpoints and logs stay in a local output folder. Optional FlashAttention and xformers support can speed up attention computation.

Alongside the downloadable code and weights, a hosted browser demo provides interactive text-to-image generation. Infinity is a separate related project for text-to-image research; VAR's repository supplies the image generation models and training implementation described here.

Similar to VAR