
OpenDataLoader PDF is a local PDF parser for developers building document search and RAG systems, with free accessibility auditing and auto-tagging. It converts PDFs into Markdown, JSON and HTML while preserving the structure needed to retrieve passages and cite their source. The open-source parser uses the Apache 2.0 license and runs locally without a GPU or API keys.
Its layout analysis handles multi-column reading order, heading hierarchy and lists. Table extraction retains row and column relationships, including merged cells. JSON output gives each element page references and bounding boxes, so an application can highlight the original passage rather than point readers to an entire document. For PDFs with reliable structure tags, it can use the author's tagged reading order and layout.
Optional hybrid mode adds AI processing for complex or borderless tables, scanned pages, LaTeX formulas and picture or chart descriptions. OCR choices include EasyOCR, RapidOCR and Tesseract, with Apple Vision available on macOS. Python, Node.js and Java SDKs, Docker support and a LangChain integration make it usable within existing document pipelines.
Accessibility checks detect untagged PDFs, and auto-tagging creates Tagged PDFs as a foundation for PDF/UA work. The parser also filters hidden text, off-page content and suspicious invisible layers that can carry prompt injection attempts.
Claim this page with an email at opendataloader.org. OpenDataLoader PDF gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find OpenDataLoader PDF?Promote it
Something wrong or outdated on this page?
9.4KUpdated 1 day agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
15.5KUpdated 4 days agoApache-2.0
macOS · Windows · Linux · Docker#Multilingual
1.6KUpdated 2 weeks agoApache-2.0
macOS · Windows · Linux · Docker#MCP#Multimodal input#Ollama integration
6.3KUpdated 3 months agoApache-2.0
Windows · Linux · Docker#Multilingual#ONNX
PaddleX is a low-code AI development toolkit for developers building document processing, computer vision and time-series applications on their own hardware. Built on PaddlePaddle, it combines pretrained models with tools for training, inference and deployment. It's open source under Apache 2.0.
68.2KUpdated 21 hours agoMIT
macOS · Windows · Linux#MCP#Works offline
40.1KUpdated 3 weeks agoApache-2.0
macOS · Linux · Web#Batch processing#llama.cpp backend#Multilingual
Unstructured is a local document processing library for developers building LLM applications and document ingestion pipelines. It turns PDFs, Word documents, HTML, emails and images into document elements that applications can use. The Python library is open source under Apache 2.0 and runs on your own hardware, including through Docker images for x86_64 and Apple Silicon.
ExtractThinker is a Python library for developers who need structured data from documents, such as invoice fields their application can use. It pairs document parsers with an LLM and returns results that follow a Pydantic schema. It's open source under Apache 2.0 and runs on macOS, Windows and Linux.
Docling is an MIT-licensed, open source document parser for developers turning files into structured content for search and AI applications. It runs locally on macOS, Linux, and Windows, including in air-gapped environments. Its PDF processing identifies page layout and reading order, extracts tables, code, and formulas, and classifies images.
Marker is a local document converter for developers and teams turning PDFs, scans and Office files into structured text. It preserves tables, equations and page structure for document processing and AI workflows. Its pipeline reads embedded PDF text and uses Surya OCR where text is missing or damaged, rather than reading every page through a vision model.