Ollama tutorial: setup, Llama 3 and Python API

Learn Ollama installation, model management and Python integration, including API requests on port 11434 and custom behavior through Modelfiles.

Player not loading? Watch on YouTube

This beginner tutorial explains how to use Ollama to run models locally on Windows, macOS and Linux. It covers installation, terminal checks and a first chat with Llama 3, then explains how to download models separately, list installed models and remove files to reclaim storage.

Hardware guidance is central to the walkthrough. The speaker presents 8 GB of RAM as a minimum for smaller models, recommends 16 GB for a smoother Llama 3 experience and suggests 32 GB or more for larger models or multitasking. These are the video's recommendations rather than guarantees for every model. It estimates 4 to 5 GB of storage for a typical 7B or 8B model and recommends starting with 20 to 50 GB free. Downloads require a connection; the speaker describes subsequent local inference as usable offline.

The programming section introduces the local server on port 11434 and the /api/generate endpoint. It explains curl requests, JSON payloads and Windows quoting, then uses Python 3.13 in the Windows setup example. Python examples cover the official Ollama package and streamed responses through requests.

A Modelfile example sets a base model, temperature and a Mario system prompt. Later sections discuss memory pressure and service troubleshooting. The closing overview introduces LLaVA image input and Open WebUI as a browser interface with chat history and file uploads, without providing a full setup walkthrough for either.