Player not loading? Watch on YouTube
This walkthrough builds a graph-based RAG question-answering system around articles about AI copyright and governance. The presenter implements the method in her own LlamaIndex notebooks, following the community detection and summarization approach described in Microsoft Research's GraphRAG paper. She recommends graph-based retrieval for questions about relationships and themes across documents, and vector retrieval for direct fact lookups. The tutorial does not run a controlled comparison between the two methods.
SerpApi sponsors the video and supplies Google search results for the collection notebook. Trafilatura extracts article text, and YouTube Transcript API retrieves captions. The notebook removes duplicate URLs, keeps successful extractions and exports a CSV.
The graph pipeline uses LlamaIndex's property graph index with a custom extractor and graph store. It requires an OpenAI API key: GPT-4o mini extracts entities and relationships, generates community summaries and answers questions against individual summaries; GPT-4o combines the relevant partial answers. The configuration allows up to 50 articles, 20 relationship triplets per article and four extraction workers. The presenter skips chunking because these articles fit within the model's context window.
An ontology and Pydantic schemas constrain extraction, while hierarchical Leiden groups entities into communities. A D3.js graph helps inspect connections. The demo returns answers about legal themes and company relationships, but a policy comparison receives an insufficient-information response. Inconsistent entity names also need normalization. These examples use hosted models and reflect the collected dataset's limits.