Player not loading? Watch on YouTube
This video explains the footage crawler built for Van Neistat's 26-year archive. Neistat reports 451 hours of footage across roughly 48,000 clips, totaling 9 TB. The workflow creates text descriptions linked to video filenames so searches can use a small index instead of analyzing the footage again for each query. He says the resulting descriptions occupy 12 MB.
FFmpeg extracts selected frames: the 50th frame, a frame three seconds before the end, and frames every 30 seconds between them. Neistat says most of his clips last less than a minute, so each usually needs only a few images. This sampling approach describes selected moments rather than every frame in a clip.
Claude guides him toward Ollama to host Moondream, the local vision model that describes the images. According to Neistat, the labeling ran on a Mac mini for two days of constant processing. The local labeling process avoids per-request cloud inference charges.
A search for "dog" returns a time-lapse clip. Neistat initially suspects a mistaken label, then spots a dog briefly appearing in the footage. The clip illustrates finding short moments through searchable labels; software setup is described at a high level.