
LayoutParser is an open-source Python library for developers and researchers who need to detect page structure in document images and turn OCR output into structured data. Its pretrained deep learning models share a common interface, so you can work with models trained on different document datasets without rewriting the surrounding pipeline.
The library covers both detection and the work that follows it. You can select text or layout elements within a page region, such as a column, and perform OCR on individual detected regions. Table OCR examples show how to parse recognition results into structured output.
Layout data has explicit representations for coordinates, text blocks, and whole layouts, with shared operations for manipulating them. This gives document processing code a consistent way to handle detected regions and their text. Visualization tools can overlay boxes and categories on the source image or redraw recognized text at its page positions, which helps you inspect detection and OCR results.
LayoutParser loads layout data from JSON, CSV, and PDFs, and exports it to JSON and CSV for use in other processing pipelines. For documents that need a model trained on their particular structure, the associated training tools support training from scratch or fine-tuning an existing model. Annotation tools use model predictions as a starting point and help reviewers find missed objects and correct labels.
The library uses the Apache 2.0 license.
Claim this page and we'll verify you by hand. LayoutParser gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find LayoutParser?Promote it
Something wrong or outdated on this page?
7.1KUpdated 2 days agoApache-2.0
Docker#Batch processing#Distributed execution#Multimodal input
Data-Juicer is a Python framework for preparing AI datasets on your own machine or a distributed Ray cluster. It's for researchers and teams curating model training data, agent interaction records or documents for retrieval. The project is open source under Apache 2.0.
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
68.2KUpdated 1 day agoMIT
macOS · Windows · Linux#MCP#Works offline
6.4KUpdated 1 day agoApache-2.0
Docker · Web
docTR is an open-source Python OCR library for developers building document processing tools and researchers comparing text recognition models. It reads PDFs and images on your own hardware, locating words and recognizing their text. The library uses PyTorch and carries the Apache 2.0 license.
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
40.1KUpdated 2 weeks agoApache-2.0
macOS · Linux · Web#Batch processing#llama.cpp backend#Multilingual
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
Docling is an MIT-licensed, open source document parser for developers turning files into structured content for search and AI applications. It runs locally on macOS, Linux, and Windows, including in air-gapped environments. Its PDF processing identifies page layout and reading order, extracts tables, code, and formulas, and classifies images.
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
Marker is a local document converter for developers and teams turning PDFs, scans and Office files into structured text. It preserves tables, equations and page structure for document processing and AI workflows. Its pipeline reads embedded PDF text and uses Surya OCR where text is missing or damaged, rather than reading every page through a vision model.