PDF-Extract-Kit is a local AI model toolbox for developers and researchers building document processing applications. It extracts text, tables and mathematical formulas from PDFs, with separate models for identifying page elements and recognizing their contents. It's open source under AGPL-3.0, written in Python, and supports CPU or GPU execution on your own hardware.
Its modular design lets developers combine the parts they need for document Q&A, translation or document assistants. Layout detection uses DocLayout-YOLO, YOLO-v10 or LayoutLMv3 to locate text, titles, images, tables and formulas. PaddleOCR reads text from images, while YOLOv8 detects inline and block equations.
Formula recognition uses UniMERNet to turn equation images into LaTeX, including long expressions, handwriting and noisy screenshots. For tables, PaddleOCR with TableMaster and StructEqTable provide recognition options. StructEqTable uses InternVL2-1B and can output LaTeX, HTML or Markdown, with support for Chinese table content.
The models use fine-tuning data drawn from different document types, including academic papers, textbooks and financial reports, rather than relying solely on a single academic dataset. The layout models also address blurred pages and watermarks. The toolkit focuses on extracting individual document elements; MinerU builds on these models to handle complete document conversion, including PDFs to Markdown.
Claim this page and we'll verify you by hand. PDF-Extract-Kit gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find PDF-Extract-Kit?Promote it
Something wrong or outdated on this page?
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
6.7KUpdated 2 months agoApache-2.0
Windows · Docker · Web#Batch processing#Hugging Face integration#Multilingual
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
40.1KUpdated 2 weeks agoApache-2.0
macOS · Linux · Web#Batch processing#llama.cpp backend#Multilingual
90.4KUpdated 2 weeks agoApache-2.0
Web#Multilingual#ONNX#Structured output
21.4KUpdated 3 weeks agoApache-2.0
macOS · Web#Batch processing#llama.cpp backend#Multilingual
MonkeyOCR is a local AI document parser for developers and researchers working with English and Chinese PDFs or images. It extracts text, formulas and tables while identifying page structure and relationships between blocks. That makes it useful for documents where plain text extraction loses reading order or separates content from its layout.
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
Marker is a local document converter for developers and teams turning PDFs, scans and Office files into structured text. It preserves tables, equations and page structure for document processing and AI workflows. Its pipeline reads embedded PDF text and uses Surya OCR where text is missing or damaged, rather than reading every page through a vision model.
PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.
Surya is a local OCR toolkit for developers extracting text and structure from PDFs and document images. It combines text recognition, layout analysis and table recognition in one vision-language model, so results retain page structure and reading order rather than just the words.