Subtitle Edit 5: Whisper and Crisp ASR setup

Learn to download speech-to-text engines and models in Subtitle Edit 5, choose CPU builds, and recognize the install-status issue shown in the tutorial.

Player not loading? Watch on YouTube

This tutorial explains how to download speech-to-text engines and models in Subtitle Edit 5 and beyond. The presenter works on Windows and says he believes the process is similar on Mac and Linux; he does not demonstrate those platforms.

After loading a video, he opens Video > Speech to text, selects the Whisper Const-me engine, and downloads its multilingual base model. Green check marks indicate that the engine and model are up to date. He then starts a local transcription and shows options for translation to English, post-processing, and advanced settings. The sample video is short, so the demonstration does not establish performance on longer recordings.

The Crisp ASR setup covers engine builds, models, and aligners. The presenter chooses the standard CPU build, described in the interface as recommended for most machines. The legacy build is a fallback for older CPUs without AVX2 support. He selects a 489 MB model and recommends downloading the required components before starting transcription, with disk space in mind.

The final section shows an unresolved installation-status issue. Vulkan, CUDA, and later CPU selections report an unknown status or no install record. The presenter suggests reporting problems on the project's GitHub page and questions whether switching builds requires another download. These observations describe his session rather than a confirmed failure on every installation.