Favicon of Qwen2.5-VL

Qwen2.5-VL

A local vision-language model for image analysis, document extraction and video understanding, with Apache 2.0 weights and Hugging Face Transformers support.

Screenshot of Qwen2.5-VL website

Qwen2.5-VL is a vision-language model you can run on your own hardware to answer questions about images and video. It's aimed at developers building document processing tools, visual assistants and agents that interact with computer or phone screens. The instruction-tuned 7B model has Apache 2.0 licensing and works with Hugging Face Transformers, with weights available in Safetensors format.

Its document capabilities go beyond reading text. It can interpret charts, graphics and page layouts, then return structured content from scanned invoices, forms and tables. It also identifies objects and reports their positions as bounding boxes or points, with JSON output for coordinates and attributes. Multiple images can appear in one query, so an application can ask it to compare pictures or reason across them.

For video, the model can analyze footage longer than an hour and locate the segments associated with a particular event. It accounts for frame timing to connect its answers to specific moments. Its visual agent capabilities include reasoning about screens and directing tools for computer and phone interaction.

Local image and video files can feed local inference, so those inputs don't need a cloud inference service. The supplied Transformers examples use CUDA GPUs. Image resolution is adjustable to balance visual detail against computation and memory use, and FlashAttention 2 can reduce memory use for workloads involving multiple images or video.

Similar to Qwen2.5-VL