Sokuji setup: local speech translation on Windows

Learn to configure Sokuji's three local models, select supported languages, and reduce translation delay with streaming models and WebGPU.

Player not loading? Watch on YouTube

The original video is in Chinese. This tutorial walks through Sokuji, an open source speech translation app described as supporting Windows, Linux, and macOS. The demonstration uses a Windows installer and configures speech recognition, machine translation, and text-to-speech to run models locally. The speaker says this setup uses no online APIs.

For Windows, the presenter lists Windows 10 or 11, at least 4 GB of RAM, and over 200 MB of disk space. Model downloads need additional storage: the selected Translate Gemma model is described as roughly 3 GB, and the speech synthesis model as 380 MB. A GPU is optional; the speaker recommends at least 8 GB of VRAM when using an Nvidia card.

The walkthrough covers model downloads, microphone selection, source and output languages, and enlarged subtitles. It changes speech synthesis speed from 1.0 to 1.2 and explains the silence and minimum speech duration settings. The presenter notes recognition errors in the demonstration and suggests their accent may contribute.

Language compatibility is a practical limitation. The speaker warns that incompatible language support across the three models can break the translation chain. For lower delay, they recommend streaming models and WebGPU where available. Optional online providers and Sokuji's cloud login are also discussed; the presenter says the cloud route sends data through its servers, whereas the demonstrated local setup keeps data on the computer.