Claude Code + RAG-Anything: LightRAG setup tutorial

Learn to add scanned PDFs and charts to LightRAG with RAG-Anything, local MinerU parsing, and a Claude Code skill for Python-based ingestion.

Player not loading? Watch on YouTube

This tutorial adds RAG-Anything to an existing self-hosted LightRAG setup and uses Claude Code to manage document ingestion and queries. It assumes a running Docker LightRAG instance and familiarity with its knowledge graph.

The speaker explains how MinerU separates scanned documents into text blocks, charts, equations and images. Local parsing models extract text, while image regions follow a separate model-processing path. In his explanation, the resulting embeddings, entities and relationships merge into the LightRAG store so users can query both text and visual material through the same API.

The demonstrated configuration uses GPT-5.4-Nano and text-embedding-3-large through OpenAI. The speaker mentions Ollama as an option for keeping model processing local, but does not demonstrate that setup. The cloud-backed example should therefore not be treated as an offline workflow.

Setup includes matching storage paths to the existing container and correcting an embedding double-wrap issue identified by the speaker. In this workflow, text uploads still use the LightRAG interface, but non-text ingestion requires a Python script wrapped in a Claude Code skill. That skill also restarts the Docker container. The speaker says parsing defaults to CPU and GPU use requires a different PyTorch version. The demo queries monthly revenue from a chart in a fictional company PDF and returns a monthly breakdown.