Favicon of Marker

Marker

Local document converter turns PDFs and Office files into Markdown, JSON or HTML, with OCR on CPU, NVIDIA GPUs or Apple Silicon and optional LLM support.

Marker is a local document converter for developers and teams turning PDFs, scans and Office files into structured text. It preserves tables, equations and page structure for document processing and AI workflows. Its pipeline reads embedded PDF text and uses Surya OCR where text is missing or damaged, rather than reading every page through a vision model.

Inputs include images, PPTX, DOCX, XLSX, HTML and EPUB alongside PDFs. Outputs include Markdown, HTML, JSON and chunks. Marker saves extracted images, removes recurring headers and footers, and retains links, references and code blocks. Equations become LaTeX; JSON keeps page and block structure. It also supports multilingual OCR, batch conversion and custom formatting logic.

Conversion runs on your hardware. Marker uses vLLM with NVIDIA GPUs or llama.cpp on CPU and Apple Silicon. Balanced mode favors conversion quality, while fast mode reduces vision model use. Text-only extraction skips OCR, so it doesn't handle scanned pages or equations.

Optional LLM processing can merge tables across pages and extract form values. Ollama supports local processing; Gemini, Claude and OpenAI-compatible services are also supported, with document content sent to the chosen service. Datalab's separate hosted platform uses Chandra and processes documents on managed infrastructure.

The Python code uses Apache 2.0. Model weights have a separate modified AI Pubs Open Rail-M license with restrictions on commercial use; check those terms for your deployment.

Similar to Marker