Player not loading? Watch on YouTube
Cognee is an open source memory framework for an AI agent. The video explains its extract, cognify, load pipeline: a language model identifies entities and relationships in documents, then stores them in a knowledge graph alongside vector embeddings. The speaker contrasts this approach with retrieval based on similar text chunks.
The local LLM demo uses Ollama with Llama 3.1 and Nomic Embed Text for embeddings. Three facts connect a robot called Magpie, its controller and a firmware recall. The presenter reports a graph with 19 nodes and 19 edges, then asks a question that requires connections across documents. The built-in visualizer shows those relationships. The description identifies the demo GPU as an RTX 4060.
According to the speaker, Cognee defaults to SQLite, LanceDB and an embedded graph. Version 1.0 supports permanent graph memory plus conversation memory that syncs into the graph, and can consolidate the memory layer into one Postgres instance. The video also discusses an MCP server and a Claude Code plugin for persistent coding assistant memory.
The review gives practical limits: Cognee is a framework with a learning curve, smaller local models produce weaker graphs in the speaker's assessment, and cognify can be slow because each chunk goes through the model. Both the language model and embedding configuration need attention; the speaker warns that configuring only one side can trigger an OpenAI fallback that requests an API key.