Favicon of LLaVA

LLaVA

A self-hosted vision-language model project for image chat, with a local Gradio interface, GPU inference and Apache 2.0 code.

Screenshot of LLaVA website

LLaVA is a family of vision-language models for researchers and developers who want to ask questions about images on their own hardware. It pairs a CLIP vision encoder with a language model to support image descriptions, visual reasoning and reading text in pictures. Its Python code is open source under Apache 2.0; the project places research-use restrictions on its data and checkpoints, with additional terms from the underlying models.

A local Gradio browser interface lets you chat with an image and compare model checkpoints in the same session. Model inference runs on your machine or self-hosted server. The project also links to hosted demos, which run separately from that local setup.

The original models use Vicuna, while LLaVA-NeXT includes variants based on Llama 3 and Qwen 1.5. The NeXT family also supports video tasks. Public checkpoints, visual instruction datasets and training code make the project relevant to teams studying multimodal models or adapting them for research.

GPU memory affects which models you can run. The model worker supports multiple GPUs and 4-bit or 8-bit inference to reduce memory use, potentially allowing inference with as little as 12GB of VRAM, though quantization can reduce accuracy. Apple M1 and M2 devices can use the MPS backend. SGLang provides an alternative GPU serving backend for higher throughput, but its LLaVA integration doesn't support 4-bit quantization.

Similar to LLaVA