Player not loading? Watch on YouTube
Cedric Clyburn demonstrates Docling, an open source CLI and library for turning documents into Markdown, JSON and a Pydantic document representation. He explains how flattened PDF text can lose table relationships, captions and reading order, then shows how OCR and layout analysis preserve more of that structure. He says document conversion can run locally on CPU, including in air-gapped environments.
The notebook walkthrough converts Docling's eight-page research paper, inspects page elements and exports eight tables as data frames. Further examples extract pictures with their captions and embedded text, and display bounding boxes around document elements. For image descriptions, Clyburn connects a locally running Granite vision model through Ollama's OpenAI endpoint.
A chunkless RAG example uses a Markdown document outline as its retrieval index. An LLM selects relevant sections and retrieves their text without a chunker, embedding model or vector database. The demonstration also queries an IBM annual report with 418 sections; these examples do not establish retrieval accuracy across other documents.
The final section covers a self-hosted REST API with Docling Serve and document processing through the Docling MCP server. Clyburn shows server checks and MCP configuration so an AI agent can request conversions, summaries and Markdown exports.