Open WebUI and Ollama: setup on Unraid

Set up Open WebUI with Ollama on Unraid, choose a model for your GPU memory, and query household manuals through a Knowledge collection.

Player not loading? Watch on YouTube

This tutorial builds a self-hosted ChatGPT alternative on an Unraid server using Open WebUI and Ollama. Open WebUI supplies the chat interface; Ollama loads and runs the models. The presenter describes the demonstrated configuration as private and leaves the optional OpenAI cloud connection unused.

The setup starts with an Ollama container suited to the graphics card. The presenter explains the Nvidia and AMD branches, their required Unraid plugins, and a separate Intel GPU container and driver plugin. CPU operation is also an option, using system memory, though he says it is slower. He covers model storage and the template's five-minute keep-alive setting. His reported 13-watt idle reading comes from his own 5090, rather than a general power guarantee.

For Open WebUI, the walkthrough changes an occupied port to 8083, generates a web secret key, enables the Ollama API and uses the server's local IP address as the connection target. A local admin account completes the initial setup.

Model selection focuses on parameter count, Q4 quantization and available VRAM. The presenter suggests Qwen 3 14B for a 12 GB card and downloads Gemma 3 27B for his demonstration. He advises leaving memory headroom for longer chats and document searches. The final examples upload household manuals to a Knowledge collection, select it with # and inspect a cited answer. A vision model also answers questions about a meal photo.