PDF and Document Parsers for RAG

Converters like Docling, MinerU and Marker turn PDFs, Office files and scans into clean Markdown or into chunks ready for RAG.

27 tools
Favicon of LightRAG

LightRAG

3 videos
An open-source RAG framework under the MIT license. Run it locally or in Docker with local models or hosted LLMs to query your documents.

39.9KUpdated 4 days agoMIT

macOS · Windows · Linux · Docker · Web#Knowledge graphs#LLM tracing#Multimodal input

An open-source RAG engine for document search and AI agents, with self-hosted deployment, hybrid retrieval, citations and a separate cloud service.

91.5KUpdated 7 hours agoApache-2.0

macOS · Windows · Linux · Docker#Hybrid search#MCP#Multi-agent workflows

Favicon of Docling

Docling

2 videos
Open source document parser that runs locally on macOS, Linux, and Windows, with OCR and structured exports for AI applications.

68.2KUpdated 1 day agoMIT

macOS · Windows · Linux#MCP#Works offline

Favicon of MinerU

MinerU

1 video
Local document parser converts PDFs, Office files and images into Markdown, JSON and LaTeX on macOS, Windows and Linux, with optional remote parsing.

80.8KUpdated 1 day ago

macOS · Windows · Linux#llama.cpp backend#MCP#MLX

Self-hosted AI application platform for enterprise workflows, RAG and agents. Runs through Docker with a browser interface under Apache 2.0.

12KUpdated 1 week agoApache-2.0

Docker · Web#Batch processing#Human approval#Multi-agent workflows

Local AI PDF parser that converts scientific papers to Markdown with LaTeX math and tables. Runs on CPU or GPU, with MIT code and CC-BY-NC weights.

10.1KUpdated 2 years agoMIT

Windows#Batch processing

Self-hosted document Q&A system with offline support, Chinese and English retrieval, and Ollama or OpenAI-compatible model connections. Licensed under AGPL-3.0.

14.2KUpdated 1 week agoAGPL-3.0

macOS · Windows · Linux · Docker · Web#Multilingual#Ollama integration#OpenAI-compatible API

An open-source Python document parser for LLM applications. Run it locally or in Docker to process PDFs, Word files and images under Apache 2.0.

15.5KUpdated 3 days agoApache-2.0

macOS · Windows · Linux · Docker#Multilingual

A self-hosted multimodal retrieval engine for AI apps, with ColPali visual search, Docker deployment, and MCP access through Claude or Open Web UI.

3.7KUpdated 5 days ago

Docker · Web#MCP#Multi-user access#Multimodal input

A self-hosted AI retrieval system with document search, cited answers and research agents. Runs in Python or Docker and uses the MIT license.

8KUpdated 11 months agoMIT

Docker#Hybrid search#Knowledge graphs#Multi-user access

An open-source Python OCR library that reads PDFs and images locally with PyTorch, runs on CPU or GPU, and supports Docker deployment.

6.4KUpdated 1 day agoApache-2.0

Docker · Web

An open-source document extraction engine with a Rust core, CPU-only processing, Docker deployment, and support for Ollama, LM Studio, and vLLM.

9.4KUpdated 2 days agoMIT

macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP

Open-source OCR library for Node.js and Python that converts documents to Markdown using cloud vision models. MIT licensed; page images leave your machine.

12.3KUpdated 1 year agoMIT

Linux#Multimodal input#Structured output

Local OCR model that converts images and PDFs to text or Markdown. Runs on NVIDIA GPUs with vLLM or Transformers under the MIT license.

23.9KUpdated 8 months agoMIT

Linux#Batch processing#Hugging Face integration#Multimodal input

A local PDF extraction toolkit in Python with OCR, layout detection and table recognition. Runs on CPU or GPU and uses the AGPL-3.0 license.

10KUpdated 2 years agoAGPL-3.0

#Hugging Face integration#Multilingual

Local OCR and document parsing software that converts PDFs and images to Markdown, recognizes tables and formulas, and runs on NVIDIA GPUs.

6.7KUpdated 2 months agoApache-2.0

Windows · Docker · Web#Batch processing#Hugging Face integration#Multilingual

Open-source AI data processing framework for local machines and Ray clusters, with multimodal cleaning, deduplication and Apache 2.0 licensing.

7.1KUpdated 2 days agoApache-2.0

Docker#Batch processing#Distributed execution#Multimodal input

Open-source Python toolkit for document layout detection and OCR workflows, with pretrained deep learning models and an Apache 2.0 license.

5.8KUpdated 4 years agoApache-2.0

An open-source codebase digest tool for LLM prompts. Process local directories with Python or the CLI, self-host with Docker, or use the hosted website.

15.7KUpdated 1 year agoMIT

Docker · Web

Favicon of Crawl4AI

Crawl4AI

1 video
Open-source web crawler that runs in Python or Docker, converts pages to Markdown and JSON, and offers a hosted API with MCP access.

84.5KUpdated 6 days agoApache-2.0

Docker#Structured output

Self-hosted web extraction API converts pages and PDFs to Markdown for LLMs. Apache 2.0 service code runs in Docker; a hosted API is also available.

12.1KUpdated 4 months agoApache-2.0

Docker#Multimodal input#Structured output

Local document converter turns PDFs and Office files into Markdown, JSON or HTML, with OCR on CPU, NVIDIA GPUs or Apple Silicon and optional LLM support.

40.1KUpdated 2 weeks agoApache-2.0

macOS · Linux · Web#Batch processing#llama.cpp backend#Multilingual

Favicon of LlamaIndex

LlamaIndex

2 videos
An MIT-licensed Python framework for document-based AI apps. It works with Ollama, while its separate document services run locally or in the cloud.

52.4KUpdated 2 days agoMIT

#Ollama integration#RAG#Reranking

An open-source OCR toolkit that converts PDFs and images into Markdown using a local GPU or an OpenAI-compatible inference server. Apache 2.0 licensed.

19.7KUpdated 6 months agoApache-2.0

Linux · Docker · Web#Batch processing#Distributed execution#OpenAI-compatible API

A self-hosted OCR model that extracts text and page structure from PDFs and images, with vLLM, Hugging Face Transformers and CPU inference support.

9.2KUpdated 6 months agoMIT

Docker#Hugging Face integration#Multilingual#Multimodal input

Local OCR software for PDFs and images, with reading order, tables and math. Runs on CPU, Apple Silicon or NVIDIA GPUs; code uses Apache 2.0.

21.4KUpdated 3 weeks agoApache-2.0

macOS · Web#Batch processing#llama.cpp backend#Multilingual

Favicon of PaddleOCR

PaddleOCR

2 videos
An open source OCR toolkit that runs locally, converts PDFs and images to Markdown or JSON, and supports multilingual text under Apache 2.0.

90.4KUpdated 2 weeks agoApache-2.0

Web#Multilingual#ONNX#Structured output

More in RAG and Vector Search