PageIndex tutorial: vectorless RAG vs vector RAG

Learn to build a PageIndex document tree with GPT-4.1 and compare retrieval against a vector RAG pipeline using two-page chunks.

Player not loading? Watch on YouTube

This tutorial compares PageIndex's reasoning-based retrieval with a conventional vector RAG pipeline on the DeepSeek R1 research paper. The speaker explains how fixed chunks can separate answers from their references, and how similarity search can retrieve a related section that lacks the requested detail.

PageIndex builds a document tree with section summaries and page locations. An LLM selects relevant nodes, then the pipeline reads the corresponding PDF pages to answer the question. In the demonstration, the paper has no table of contents, so PageIndex constructs one from its headings and saves the structure as JSON.

The setup covers cloning the repository, installing dependencies, configuring endpoints in a .env file and authenticating through az login. It uses Microsoft Foundry and Azure OpenAI with GPT-4.1. Although the repository is cloned for a self-hosted setup, the demonstrated model calls use cloud services; this is not an offline workflow.

The comparison uses 43 vector chunks, each spanning two pages, against 63 semantic nodes. Four questions test retrieval across section boundaries and terminology differences. The speaker reports more targeted page selection with PageIndex, but also notes slower retrieval and an answer with unnecessary detail. These results concern one document and a particular baseline. The speaker recommends vector retrieval for many small, independent documents and tree navigation for long documents with cross-references.