
Qwen2.5-VL is a vision-language model you can run on your own hardware to answer questions about images and video. It's aimed at developers building document processing tools, visual assistants and agents that interact with computer or phone screens. The instruction-tuned 7B model has Apache 2.0 licensing and works with Hugging Face Transformers, with weights available in Safetensors format.
Its document capabilities go beyond reading text. It can interpret charts, graphics and page layouts, then return structured content from scanned invoices, forms and tables. It also identifies objects and reports their positions as bounding boxes or points, with JSON output for coordinates and attributes. Multiple images can appear in one query, so an application can ask it to compare pictures or reason across them.
For video, the model can analyze footage longer than an hour and locate the segments associated with a particular event. It accounts for frame timing to connect its answers to specific moments. Its visual agent capabilities include reasoning about screens and directing tools for computer and phone interaction.
Local image and video files can feed local inference, so those inputs don't need a cloud inference service. The supplied Transformers examples use CUDA GPUs. Image resolution is adjustable to balance visual detail against computation and memory use, and FlashAttention 2 can reduce memory use for workloads involving multiple images or video.
Claim this page and we'll verify you by hand. Qwen2.5-VL gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Qwen2.5-VL?Promote it
Something wrong or outdated on this page?
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
huggingface.coComputer Vision Models
#Hugging Face integration#Multimodal input#Structured output
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
8.2KUpdated 2 years ago
#Batch processing#GGUF#Hugging Face integration
GOT-OCR2.0 is an OCR model for developers and researchers who want to extract text from images on their own hardware. It handles both plain text and formatted output through a single model, with recognition modes for selected regions and documents spanning multiple pages. The Python codebase builds on Vary.
21.4KUpdated 3 weeks agoApache-2.0
macOS · Web#Batch processing#llama.cpp backend#Multilingual
90.4KUpdated 2 weeks agoApache-2.0
Web#Multilingual#ONNX#Structured output
Florence-2 is Microsoft's open-source vision model for developers who want to process images on their own hardware. It handles several image tasks through text prompts, so one model can generate descriptions, read text and locate objects. It runs locally with PyTorch and Hugging Face Transformers on a CPU or CUDA GPU, and uses the MIT license.
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
Surya is a local OCR toolkit for developers extracting text and structure from PDFs and document images. It combines text recognition, layout analysis and table recognition in one vision-language model, so results retain page structure and reading order rather than just the words.
PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.