Favicon of PageIndex

PageIndex

Open-source document RAG engine with local indexing and storage, page citations, and LLM retrieval over document trees. MIT licensed.

Screenshot of PageIndex website

PageIndex is an open-source document RAG engine for developers building question answering over long PDFs, such as financial reports, legal documents and technical manuals. It organizes documents into a hierarchical tree and uses an LLM to find relevant sections, rather than relying on vector similarity search. It doesn't require a vector database or document chunking.

The Python SDK supports local indexing, retrieval and chat with your own LLM key. Document indexing and storage stay on your machine in local mode; model calls use the LLM you connect, so local mode doesn't itself mean offline inference. The MIT-licensed version focuses on text-based PDFs and provides page-level citations. Retrieval can take conversation history and domain knowledge into account, and it reads selected sections instead of sending the entire PDF for every question.

For applications, the SDK supports streaming responses and searches across multiple documents. You can connect its retrieval tools to the OpenAI Agents SDK or Claude Agent SDK. Its explicit document references help readers check answers against the original pages.

PageIndex Cloud runs document indexing and storage on PageIndex's servers. It adds OCR and image understanding for scanned or image-rich documents, block-level citations, and an MCP server. Chat and retrieval remain compatible with your chosen model provider. The cloud offering also has a file-level tree index for searching across a document corpus; dedicated VPC and on-premises deployments are available.

Similar to PageIndex