Moondream, Ollama and FFmpeg: local video search workflow

Learn how frame sampling and local vision descriptions made Van Neistat's 9 TB video archive searchable, with two days of processing on a Mac mini.

Player not loading? Watch on YouTube

This video explains the footage crawler built for Van Neistat's 26-year archive. Neistat reports 451 hours of footage across roughly 48,000 clips, totaling 9 TB. The workflow creates text descriptions linked to video filenames so searches can use a small index instead of analyzing the footage again for each query. He says the resulting descriptions occupy 12 MB.

FFmpeg extracts selected frames: the 50th frame, a frame three seconds before the end, and frames every 30 seconds between them. Neistat says most of his clips last less than a minute, so each usually needs only a few images. This sampling approach describes selected moments rather than every frame in a clip.

Claude guides him toward Ollama to host Moondream, the local vision model that describes the images. According to Neistat, the labeling ran on a Mac mini for two days of constant processing. The local labeling process avoids per-request cloud inference charges.

A search for "dog" returns a time-lapse clip. Neistat initially suspects a mistaken label, then spots a dog briefly appearing in the footage. The clip illustrates finding short moments through searchable labels; software setup is described at a high level.