Nougat is a local AI PDF parser for researchers and developers who need scientific papers as usable text, including their equations and tables. It converts academic PDFs into Markdown-style documents with LaTeX notation, so the output retains structure that plain text extraction can lose.
The Python tool runs on your own machine and supports CPU or GPU processing, including GPU use on Windows. It provides a command-line interface for individual PDFs or batches and a local API for applications that need document conversion. You can process selected pages instead of a whole paper. Model checkpoints are available to download, with small and base variants.
Its output uses .mmd files, a markup format mostly compatible with Mathpix Markdown, with LaTeX tables. Markdown compatibility processing is enabled by default. This makes Nougat relevant for document-processing projects where mathematical content needs to survive conversion alongside the surrounding prose.
Document type matters. Nougat was trained on scientific papers from arXiv and PMC and works best with English papers. Other languages using Latin scripts may work, but Chinese, Russian and Japanese aren't supported. Its failure detection can also mark pages as missing on some CPUs or older GPUs.
The project builds on Donut and includes tools for preparing training datasets, fine-tuning models and evaluating results. The code is open source under MIT; the model weights carry a separate CC-BY-NC license that restricts commercial use.
Claim this page and we'll verify you by hand. Nougat gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Nougat?Promote it
Something wrong or outdated on this page?
6.7KUpdated 2 months agoApache-2.0
Windows · Docker · Web#Batch processing#Hugging Face integration#Multilingual
MonkeyOCR is a local AI document parser for developers and researchers working with English and Chinese PDFs or images. It extracts text, formulas and tables while identifying page structure and relationships between blocks. That makes it useful for documents where plain text extraction loses reading order or separates content from its layout.
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
68.2KUpdated 1 day agoMIT
macOS · Windows · Linux#MCP#Works offline
80.8KUpdated 1 day ago
macOS · Windows · Linux#llama.cpp backend#MCP#MLX
15.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker#Multilingual
Unstructured is a local document processing library for developers building LLM applications and document ingestion pipelines. It turns PDFs, Word documents, HTML, emails and images into document elements that applications can use. The Python library is open source under Apache 2.0 and runs on your own hardware, including through Docker images for x86_64 and Apple Silicon.
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
Docling is an MIT-licensed, open source document parser for developers turning files into structured content for search and AI applications. It runs locally on macOS, Linux, and Windows, including in air-gapped environments. Its PDF processing identifies page layout and reading order, extracts tables, code, and formulas, and classifies images.
MinerU parses documents locally into structured text for AI agents, RAG systems and knowledge bases. It's for people working with scanned PDFs, academic papers and Office files whose tables, formulas or page layouts need more care than plain text extraction.
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.