Player not loading? Watch on YouTube
This tutorial builds a local RAG application that answers questions about 10 fictional investigation PDFs inspired by The Lord of the Rings and The Matrix. It explains how retrieval supplies document context to a local LLM. The project uses that context rather than fine-tuning the model.
The walkthrough uses WSL, a Conda environment with Python 3.12, and Jupyter Lab. Ollama runs Qwen2.5:1.5B for answers and BGE-M3 for embeddings. LangChain connects the document loading, text splitting and retrieval steps, while FAISS stores the vectors. The demonstrated workflow runs on CPU.
The sample PDFs produce 49 pages. Splitting them into chunks of up to 500 characters with 150 characters of overlap produces 147 chunks. The speaker then saves the FAISS index and adjusts retrieval to return five chunks per question. An answer function joins their text into context and sends it alongside the question to the chat model.
The speaker contrasts an unsupported answer with one grounded in the fictional files. That example illustrates the retrieval process, rather than establishing general answer accuracy. The initial function handles one question at a time; conversation history requires additional code. Closing advice covers a system prompt, reloading the saved index and testing chunk sizes and retrieval counts. For GPU use, the speaker describes FAISS GPU with a CUDA-capable GPU and says Ollama selects an available GPU or falls back to CPU.