Favicon of GOT-OCR2.0

GOT-OCR2.0

Local OCR model that extracts plain or formatted text from images, supports multi-page recognition, and works with Hugging Face Transformers.

GOT-OCR2.0 is an OCR model for developers and researchers who want to extract text from images on their own hardware. It handles both plain text and formatted output through a single model, with recognition modes for selected regions and documents spanning multiple pages. The Python codebase builds on Vary.

You can target a bounding box or text in a particular color when you only need part of an image. Multi-crop recognition handles multiple crops, while multi-page recognition processes a collection of page images. The tool can also render formatted results as HTML for viewing in a browser.

The main implementation uses PyTorch and CUDA for local inference. Hugging Face Transformers supports the model and batch inference, and PaddleMIX also provides support. Community projects offer CPU inference, GGUF support through llama.cpp, and implementations using OpenVINO, ONNX and MNN. These are separate contributions rather than capabilities of the main Python implementation.

Hosted demos run on Hugging Face and ModelScope, using those services' GPU resources. Local inference uses downloaded model weights and image files on your machine.

For developers adapting recognition to their own documents, the project supports fine-tuning with ms-swift and post-training from the supplied GOT weights. It also includes evaluation code and benchmarks, including Fox and OneChart.

Similar to GOT-OCR2.0