Cognee + Ollama: local knowledge graph setup

Set up Cognee 1.1.2 with Ollama, Llama 3.1 8B and nomic-embed-text v1.5, then ingest five documents and query a local knowledge graph.

Player not loading? Watch on YouTube

Brandon Minnick walks through a laptop setup for Cognee with Ollama as the local LLM provider. The tutorial uses Cognee 1.1.2, Llama 3.1 8B and nomic-embed-text v1.5, with Python 3.12 listed in the source description. Cognee sponsors the video. Minnick describes the setup as keeping document processing on the machine without per-query API charges.

The walkthrough covers Ollama installation on macOS, Linux and Windows, model downloads, and an API check on port 11434. A Python environment created with uv holds the ingestion script. The .env configuration points generation requests to /v1 and embeddings to /api/embed, with 768 embedding dimensions and placeholder API keys that the speaker says Ollama ignores.

Model choice gets particular attention. Minnick explains that Cognee uses Instructor to obtain structured JSON matching a Pydantic schema. He warns that models below roughly 7B parameters can return malformed output that prevents useful extraction. The source description recommends about 16GB RAM for Llama 3.1 8B; the spoken RAM warning inconsistently names Llama 3.2. He also sets COGNEE_SKIP_CONNECTION_TEST=true to accommodate slower local responses.

The example ingests five company documents with cognee.remember() and asks an onboarding question through cognee.recall(). The returned answer draws on the documents' user-service access rules.