Favicon of PDF-Extract-Kit

PDF-Extract-Kit

A local PDF extraction toolkit in Python with OCR, layout detection and table recognition. Runs on CPU or GPU and uses the AGPL-3.0 license.

PDF-Extract-Kit is a local AI model toolbox for developers and researchers building document processing applications. It extracts text, tables and mathematical formulas from PDFs, with separate models for identifying page elements and recognizing their contents. It's open source under AGPL-3.0, written in Python, and supports CPU or GPU execution on your own hardware.

Its modular design lets developers combine the parts they need for document Q&A, translation or document assistants. Layout detection uses DocLayout-YOLO, YOLO-v10 or LayoutLMv3 to locate text, titles, images, tables and formulas. PaddleOCR reads text from images, while YOLOv8 detects inline and block equations.

Formula recognition uses UniMERNet to turn equation images into LaTeX, including long expressions, handwriting and noisy screenshots. For tables, PaddleOCR with TableMaster and StructEqTable provide recognition options. StructEqTable uses InternVL2-1B and can output LaTeX, HTML or Markdown, with support for Chinese table content.

The models use fine-tuning data drawn from different document types, including academic papers, textbooks and financial reports, rather than relying solely on a single academic dataset. The layout models also address blurred pages and watermarks. The toolkit focuses on extracting individual document elements; MinerU builds on these models to handle complete document conversion, including PDFs to Markdown.

Similar to PDF-Extract-Kit