GOT-OCR2.0 is an OCR model for developers and researchers who want to extract text from images on their own hardware. It handles both plain text and formatted output through a single model, with recognition modes for selected regions and documents spanning multiple pages. The Python codebase builds on Vary.
You can target a bounding box or text in a particular color when you only need part of an image. Multi-crop recognition handles multiple crops, while multi-page recognition processes a collection of page images. The tool can also render formatted results as HTML for viewing in a browser.
The main implementation uses PyTorch and CUDA for local inference. Hugging Face Transformers supports the model and batch inference, and PaddleMIX also provides support. Community projects offer CPU inference, GGUF support through llama.cpp, and implementations using OpenVINO, ONNX and MNN. These are separate contributions rather than capabilities of the main Python implementation.
Hosted demos run on Hugging Face and ModelScope, using those services' GPU resources. Local inference uses downloaded model weights and image files on your machine.
For developers adapting recognition to their own documents, the project supports fine-tuning with ms-swift and post-training from the supplied GOT weights. It also includes evaluation code and benchmarks, including Fox and OneChart.
Claim this page and we'll verify you by hand. GOT-OCR2.0 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find GOT-OCR2.0?Promote it
Something wrong or outdated on this page?
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
huggingface.coOCR and Document Scanning
Linux#Batch processing#Hugging Face integration#Multimodal input
21.4KUpdated 3 weeks agoApache-2.0
macOS · Web#Batch processing#llama.cpp backend#Multilingual
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
huggingface.coComputer Vision Models
#Hugging Face integration#Multimodal input#Structured output
90.4KUpdated 2 weeks agoApache-2.0
Web#Multilingual#ONNX#Structured output
Qwen2.5-VL is a vision-language model you can run on your own hardware to answer questions about images and video. It's aimed at developers building document processing tools, visual assistants and agents that interact with computer or phone screens. The instruction-tuned 7B model has Apache 2.0 licensing and works with Hugging Face Transformers, with weights available in Safetensors format.
Surya is a local OCR toolkit for developers extracting text and structure from PDFs and document images. It combines text recognition, layout analysis and table recognition in one vision-language model, so results retain page structure and reading order rather than just the words.
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
Florence-2 is Microsoft's open-source vision model for developers who want to process images on their own hardware. It handles several image tasks through text prompts, so one model can generate descriptions, read text and locate objects. It runs locally with PyTorch and Hugging Face Transformers on a CPU or CUDA GPU, and uses the MIT license.
PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.