Favicon of SenseVoice

SenseVoice

Open-source speech recognition model for Mandarin, Cantonese, English, Japanese and Korean. Runs locally on CPU or GPU under the MIT license.

SenseVoice is a local speech recognition model that adds language, emotion and sound-event tags to transcriptions. It's for developers building voice applications or analyzing recordings on their own hardware, particularly those working with Mandarin and Cantonese. The project is open source under the MIT license.

The released SenseVoiceSmall model transcribes Mandarin, Cantonese, English, Japanese and Korean and identifies the spoken language. It can label emotions such as happiness, sadness and anger, alongside sounds such as laughter, applause, coughing and background music. Sound detection also works as a standalone task, though specialized audio-event models perform better on some benchmarks.

Local deployment supports CPU or GPU inference, including Docker. A llama.cpp runtime uses GGUF models and runs as a self-contained CPU binary without Python or a GPU, with built-in speech segmentation. Other deployment paths use FunASR, ONNX or Libtorch. ModelScope and Hugging Face also provide hosted demos, separate from the local runtime.

SenseVoiceSmall uses a non-autoregressive architecture for low-latency transcription. Published benchmarks report faster inference than Whisper and recognition advantages on Chinese and Cantonese test sets. Long-recording support includes a bounded-memory path that processes audio in overlapping sections, and fine-tuning scripts support adaptation to specific audio data. Speaker labels require a composed FunASR pipeline with separate FSMN-VAD and CAM++ models; SenseVoiceSmall doesn't identify speakers on its own.

Similar to SenseVoice