
MinerU parses documents locally into structured text for AI agents, RAG systems and knowledge bases. It's for people working with scanned PDFs, academic papers and Office files whose tables, formulas or page layouts need more care than plain text extraction.
It separates text and images while preserving document context and layout hierarchy. Alongside PDF, Word, PowerPoint and Excel, it accepts images, EPUB, OpenDocument files and web formats, including HTML and MHTML archives. The online service also accepts web URLs. Outputs include Markdown and JSON, with LaTeX and MathML for mathematical expressions.
Complex tables are a particular focus: it handles rotated tables, cells that span pages and merged cells, then exports table data as CSV, HTML or Markdown. Formula parsing covers long expressions, multiple lines and nested structures. EPUB parsing follows the book's reading order and preserves internal links.
Documents stay local by default. Remote parsing is optional and sends documents to a remote service; local failures don't silently trigger an upload. MinerU runs on macOS, Windows and Linux, and can serve parsing requests through a local server. CPU-only parsing works, with an ONNX path requiring as little as 2 GB RAM. Apple Silicon and NVIDIA GPUs can accelerate model-based parsing, with heavier GPU paths requiring more memory.
Supported engines include llama.cpp, PyTorch, vLLM, lmdeploy and MLX. MCP and OpenClaw support connect it to agent workflows, and it can export results to Dify and Notion.
Claim this page with an email at mineru.net. MinerU gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find MinerU?Promote it
Something wrong or outdated on this page?
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
68.2KUpdated 1 day agoMIT
macOS · Windows · Linux#MCP#Works offline
15.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker#Multilingual
Unstructured is a local document processing library for developers building LLM applications and document ingestion pipelines. It turns PDFs, Word documents, HTML, emails and images into document elements that applications can use. The Python library is open source under Apache 2.0 and runs on your own hardware, including through Docker images for x86_64 and Apple Silicon.
37.3KUpdated 24 hours agoAGPL-3.0
macOS · Windows · Docker · Web#Batch processing#Hugging Face integration#MCP
40.1KUpdated 2 weeks agoApache-2.0
macOS · Linux · Web#Batch processing#llama.cpp backend#Multilingual
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
Docling is an MIT-licensed, open source document parser for developers turning files into structured content for search and AI applications. It runs locally on macOS, Linux, and Windows, including in air-gapped environments. Its PDF processing identifies page layout and reading order, extracts tables, code, and formulas, and classifies images.
PDFMathTranslate translates scientific PDFs while keeping their page layout, formulas, charts, contents pages and annotations. It's for researchers, students and others who need to read papers in another language without losing the relationship between the text and its figures. It produces both translated PDFs and bilingual documents for comparison with the original.
Marker is a local document converter for developers and teams turning PDFs, scans and Office files into structured text. It preserves tables, equations and page structure for document processing and AI workflows. Its pipeline reads embedded PDF text and uses Surya OCR where text is missing or damaged, rather than reading every page through a vision model.
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.