LLM Vision, Ollama & Home Assistant: camera setup

Learn camera analysis and visitor-log automations with LLM Vision and Ollama on a CPU-only VPS with four vCPUs and 16 GB RAM.

Player not loading? Watch on YouTube

This tutorial builds a self-hosted AI camera workflow with LLM Vision, Ollama and Home Assistant. The presenter uses a VPS with four vCPUs, 16 GB RAM and no GPU, installs Ollama through Docker, and loads a model introduced as Gemma 4 E2B QAT.

The setup covers installing LLM Vision through HACS and configuring Ollama as its provider. The presenter recommends think mode and a context length of 4096. Image Analyzer describes a camera snapshot and stores the result in a local timeline. Stream Analyzer uses a five-second recording with a maximum of three frames, while Video Analyzer accepts a local path to saved footage. These examples show text summaries rather than measured accuracy or processing speed.

Get Events retrieves timeline entries with date and person filters. Data Analyzer answers a people-count question and updates a manually created helper; the example returns four.

Two automations turn those actions into a visitor log. A presence sensor triggers snapshot analysis and a count update. A second automation runs at 23:59, retrieves the past 24 hours of events, asks the model for a summary, and sends it to Home Assistant notifications and a mobile phone.