Player not loading? Watch on YouTube
This tutorial installs SimpleMem on Ubuntu and connects it to an Ollama endpoint for local AI memory tests. The speaker creates a Conda environment, clones the repository, installs prerequisites, and edits an example configuration with a dummy API key, the local endpoint, and a model name. His setup uses a customized 27-billion-parameter model with an increased context length. The first smoke test downloads an embedding model before storing dialogue memories.
The speaker describes SimpleMem as compressing dialogue into atomic facts and indexing them by meaning, keywords, and metadata such as dates and entities. An LLM planner splits questions into retrieval requirements, searches the indexes, and checks whether the results cover the question. In the first example, he reports that the system retrieves a meeting's time and place and resolves "tomorrow" to November 16, 2025. He also explains an offline improvement loop that evaluates retrieval failures and proposes configuration changes with regression checks.
A larger test asks about a laptop and its use. The speaker judges the answer correct and reports recovery from errors, but flags slow memory ingestion: reaching that stage took around six minutes. He reports roughly 30 to 32 GB of memory use in his setup and notes that VRAM consumption depends on the model. These are observations from his demonstration, rather than general hardware requirements or guarantees of accuracy.