Favicon of OpenDataLoader PDF

OpenDataLoader PDF

Local PDF parser that extracts Markdown, JSON and HTML with source coordinates. Apache 2.0 licensed, with optional AI processing for complex pages.

Screenshot of OpenDataLoader PDF website

OpenDataLoader PDF is a local PDF parser for developers building document search and RAG systems, with free accessibility auditing and auto-tagging. It converts PDFs into Markdown, JSON and HTML while preserving the structure needed to retrieve passages and cite their source. The open-source parser uses the Apache 2.0 license and runs locally without a GPU or API keys.

Its layout analysis handles multi-column reading order, heading hierarchy and lists. Table extraction retains row and column relationships, including merged cells. JSON output gives each element page references and bounding boxes, so an application can highlight the original passage rather than point readers to an entire document. For PDFs with reliable structure tags, it can use the author's tagged reading order and layout.

Optional hybrid mode adds AI processing for complex or borderless tables, scanned pages, LaTeX formulas and picture or chart descriptions. OCR choices include EasyOCR, RapidOCR and Tesseract, with Apple Vision available on macOS. Python, Node.js and Java SDKs, Docker support and a LangChain integration make it usable within existing document pipelines.

Accessibility checks detect untagged PDFs, and auto-tagging creates Tagged PDFs as a foundation for PDF/UA work. The parser also filters hidden text, off-page content and suspicious invisible layers that can carry prompt injection attempts.

Similar to OpenDataLoader PDF