Favicon of LayoutParser

LayoutParser

Open-source Python toolkit for document layout detection and OCR workflows, with pretrained deep learning models and an Apache 2.0 license.

Screenshot of LayoutParser website

LayoutParser is an open-source Python library for developers and researchers who need to detect page structure in document images and turn OCR output into structured data. Its pretrained deep learning models share a common interface, so you can work with models trained on different document datasets without rewriting the surrounding pipeline.

The library covers both detection and the work that follows it. You can select text or layout elements within a page region, such as a column, and perform OCR on individual detected regions. Table OCR examples show how to parse recognition results into structured output.

Layout data has explicit representations for coordinates, text blocks, and whole layouts, with shared operations for manipulating them. This gives document processing code a consistent way to handle detected regions and their text. Visualization tools can overlay boxes and categories on the source image or redraw recognized text at its page positions, which helps you inspect detection and OCR results.

LayoutParser loads layout data from JSON, CSV, and PDFs, and exports it to JSON and CSV for use in other processing pipelines. For documents that need a model trained on their particular structure, the associated training tools support training from scratch or fine-tuning an existing model. Annotation tools use model predictions as a starting point and help reviewers find missed objects and correct labels.

The library uses the Apache 2.0 license.

Similar to LayoutParser