Favicon of GraphRAG

GraphRAG

An open-source Python RAG system under the MIT license that uses knowledge graphs and community summaries to answer questions about private datasets.

Screenshot of GraphRAG website

GraphRAG builds a knowledge graph from text so an LLM can answer questions that depend on connections across documents or themes across a whole collection. It's for developers and researchers working with private datasets, such as business documents, proprietary research, or communications.

Its main distinction is how it finds context. Vector search retrieves passages similar to a question; GraphRAG also extracts entities, relationships, and claims, then groups related entities into communities with layered summaries. These structures give the model context for questions whose answers span scattered information or require an overview of a large dataset.

The query modes serve different needs:

  • Global Search uses community summaries for questions about the collection as a whole.
  • Local Search follows relationships around particular entities, while DRIFT Search adds community context to that exploration.
  • Basic Search retrieves text through conventional vector similarity when that approach fits the question.

The project is open source under the MIT license and written in Python. Its indexing pipeline retains references to source text, and prompt tuning lets developers adapt it to their data. Indexing can be expensive, so the cost of preparing a collection is a practical consideration.

GraphRAG is a Microsoft Research demonstration rather than an officially supported Microsoft product. It's in maintenance mode, with bug fixes and dependency updates continuing, but no new features or outside pull requests.

Similar to GraphRAG