Player not loading? Watch on YouTube
This Russian-language tutorial is summarized from its complete original Russian captions. Uploaded on October 22, 2024, this tutorial demonstrates a Windows setup using Ollama with llama3.2:3b-instruct-fp16. The supplied description names LangChain, LangGraph and intfloat/multilingual-e5-large embeddings from Hugging Face. These are the tutorial's model and framework choices; it does not specify framework versions.
The walkthrough separates the code into three modules. The first checks that Python can reach Ollama and obtain a response from the local LLM. The second builds a RAG workflow over PDF documents: it creates a vector knowledge base, retrieves relevant chunks and asks the model to answer using them. PDFs must have a text layer. The description names E5 Embeddings’ multilingual-e5-large model, and the linked source code configures that exact model for PDF vectorization with HuggingFaceEmbeddings. The first run downloads the embedding model and builds the index; later runs reuse the saved database. The speaker explains that several PDFs can go into the source folder.
The third module adds an AI agent workflow in LangGraph. It routes questions, grades retrieved passages and falls back to web search when the documents do not supply suitable material. It also checks generated answers for hallucinations and can repeat parts of the process. Those checks are part of the demonstrated workflow, not proof that its answers are correct.
Although the speaker calls the assistant fully local, the example uses Tavily for internet searches and requires an account and API key. Model inference and document indexing run locally, but the search path is not offline.