Player not loading? Watch on YouTube
This tutorial sets up Ollama on a Debian server with an Nvidia GTX 1660 Ti, then compares responses with and without GPU acceleration. The server detects both the Nvidia card and an integrated AMD Radeon 680M. The walkthrough focuses on the Nvidia card for running a local LLM.
The setup covers adding repositories, installing the Nvidia Container Toolkit, configuring Docker and restarting it. The speaker creates a Docker Compose service for Ollama using the latest image, a persistent data folder and a GPU reservation in the deploy block. The first container launch fails. Troubleshooting includes enabling Debian's non-free and non-free-firmware repositories, installing the Nvidia driver and checking Nvidia-SMI. After another server reboot, the container starts successfully.
The speaker downloads a small Gemma model and tests a greeting and a request for help configuring Ollama with a GPU. Disabling the GPU reservation lets him compare CPU inference with the accelerated run. He reports slower text generation on the CPU, though it performs better than he expected. A larger model responds slowly, and he notes that models which do not fit in the card's memory can remain slow with GPU acceleration. These are informal observations from his server, without numerical benchmark results.