Favicon of docTR

docTR

An open-source Python OCR library that reads PDFs and images locally with PyTorch, runs on CPU or GPU, and supports Docker deployment.

Screenshot of docTR website

docTR is an open-source Python OCR library for developers building document processing tools and researchers comparing text recognition models. It reads PDFs and images on your own hardware, locating words and recognizing their text. The library uses PyTorch and carries the Apache 2.0 license.

Its pretrained models let you build an OCR pipeline without training everything yourself. You can choose detection and recognition models separately, pairing detectors such as DBNet or FAST with recognizers such as CRNN, PARSeq or ViTSTR. Researchers can also train their own models and compare them on supported public datasets, including FUNSD, CORD and SROIE.

Beyond word recognition, docTR detects document regions such as tables, figures and headers. It handles rotated pages and words, and can return rotated bounding boxes. A key information extraction pipeline combines recognition with a detector trained to find specific classes, such as dates or addresses. Predictions can be inspected visually.

Inference runs on CPU or GPU, and Docker support includes NVIDIA GPU acceleration. For teams building a self-hosted document service, the project provides a FastAPI template. A Streamlit demo runs locally in a browser; a separate demo on Hugging Face Spaces runs on hosted infrastructure.

Similar to docTR