Ollama Python: install the SDK and call local chat models

Learn to install Ollama, download Llama 3.2, and call chat() from Python. The tutorial specifies Python 3.8+ and 2.0 GB of model disk space.

Player not loading? Watch on YouTube

This tutorial introduces Ollama as an open source platform for running a local LLM and connecting it to Python. The speaker describes privacy, independence from vendor accounts, and offline use as benefits of local execution. The setup starts with installing Ollama and checking its version, then downloading Llama 3.2 from the model library. The instructor specifies 2.0 GB of disk space for the model used in the course.

The server must be running before Python can connect. On some Linux distributions, the instructor says to start it with ollama serve; on Mac or Windows, opening the desktop application may resolve a command-line connection problem. She recommends an optional virtual environment and demonstrates Python 3.14 on a Mac, while stating that the course requires Python 3.8 or higher. The Python client installs with pip install ollama.

The example calls chat() with a model name and a messages list, then prints the response's message content. A user message asks for a one-sentence description of Python. This initial script has no conversation history: earlier exchanges must be included in the list, in chronological order. Adding a system message changes the response to pirate speech, and inspecting the returned role shows assistant. The excerpt stops short of building a full chat application.