Docling: chunkless RAG vs. vector retrieval

Learn how Docling supports document navigation for chunkless RAG, and why the speaker says it needs more model calls than vector lookup.

Player not loading? Watch on YouTube

Ming Zhao explains chunkless RAG through a question about a 200-page annual report: what changed in a revenue recognition policy, and where does the report explain why? Conventional retrieval splits documents into chunks, embeds them as vectors, and retrieves text similar to the question. Zhao argues that this can separate headings, tables and explanations that belong together.

The alternative keeps a document tree. An AI agent starts with an outline and section summaries, chooses a likely section, reads it, and follows references or visits another branch when it needs more evidence. Zhao says this preserves the section context and helps answer questions whose supporting material spans several parts of a document.

Docling supplies the structure in this explanation. It reconstructs a PDF as a Docling document with headings, reading order and tables. The speaker also describes a Docling agent that can edit content, extract fields and enrich sections.

The tradeoff is extra work. Clean document parsing is difficult, and navigation requires more model calls and latency than a single vector lookup. Zhao recommends structure-aware retrieval for long, organized documents where connections matter. For broad searches across millions of documents, he favors similarity search, or a combination that finds a document first and then navigates its structure.