MemOS 2.0 setup with Qwen2.5 and Qdrant on RTX 4090

Learn to configure MemOS memory extraction and recall with Qwen2.5-3B, Qdrant and a Gradio demo on an RTX 4090, including dependency pitfalls.

Player not loading? Watch on YouTube

This walkthrough sets up MemOS 2.0, an open source memory system for language models and agents, on an RTX 4090 workstation. The presenter uses Qwen2.5-3B-Instruct for chat and memory extraction, all-MiniLM-L6-v2 for embeddings, and Qdrant for storage. The demonstrated local AI stack runs without a cloud LLM or API keys.

Setup covers Python 3.12.8, uv 0.8.8 and CUDA 12.8 PyTorch wheels. The presenter installs CUDA-enabled PyTorch before MemOS and adds accelerate separately because the Hugging Face backend needs it. A singleton configuration shares one model across chat, the memory reader and the extractor. The tutorial focuses on textual memory; it leaves activation, parametric and preference memory uninitialized.

The smoke test extracts facts from a conversation, embeds and stores them, retrieves five memories, and supplies them to the chat model. The presenter reports extraction under two seconds and a 10-millisecond search on this machine. A custom Gradio interface adds tabs for memory entry, search and chat.

The final sections discuss personal assistants, support history, shared agent memory and Python library versus REST deployment. Practical caveats include calling extract() explicitly for summaries because general-text add() stores raw text. The presenter also reports a sentence-chunker failure with the current chunky release that affects document ingestion.