SenseVoice is a local speech recognition model that adds language, emotion and sound-event tags to transcriptions. It's for developers building voice applications or analyzing recordings on their own hardware, particularly those working with Mandarin and Cantonese. The project is open source under the MIT license.
The released SenseVoiceSmall model transcribes Mandarin, Cantonese, English, Japanese and Korean and identifies the spoken language. It can label emotions such as happiness, sadness and anger, alongside sounds such as laughter, applause, coughing and background music. Sound detection also works as a standalone task, though specialized audio-event models perform better on some benchmarks.
Local deployment supports CPU or GPU inference, including Docker. A llama.cpp runtime uses GGUF models and runs as a self-contained CPU binary without Python or a GPU, with built-in speech segmentation. Other deployment paths use FunASR, ONNX or Libtorch. ModelScope and Hugging Face also provide hosted demos, separate from the local runtime.
SenseVoiceSmall uses a non-autoregressive architecture for low-latency transcription. Published benchmarks report faster inference than Whisper and recognition advantages on Chinese and Cantonese test sets. Long-recording support includes a bounded-memory path that processes audio in overlapping sections, and fine-tuning scripts support adaptation to specific audio data. Speaker labels require a composed FunASR pipeline with separate FSMN-VAD and CAM++ models; SenseVoiceSmall doesn't identify speakers on its own.
Claim this page and we'll verify you by hand. SenseVoice gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find SenseVoice?Promote it
Something wrong or outdated on this page?
18.5KUpdated 23 hours agoApache-2.0
Linux · Docker#Batch processing#Hugging Face integration
Parakeet is NVIDIA's speech recognition model family. The linked parakeet-tdt-0.6b-v2 is its English speech-to-text model for developers and researchers building transcription services, subtitles or voice applications. It runs locally through NeMo on Linux, with NVIDIA GPUs recommended for inference. It's a model you can embed in an application, rather than a desktop transcription app.
54.5KUpdated 4 weeks agoMIT
#Hugging Face integration#Multilingual#Quantization
10.8KUpdated 3 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
4.1KUpdated 2 years agoMIT
#Batch processing#Hugging Face integration
109.8KUpdated 4 weeks agoMIT
#Multilingual#Voice activity detection
11.2KUpdated 1 month ago
macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference
VibeVoice is a family of MIT-licensed, open-source voice AI models for developers and researchers building local transcription or speech generation tools. Its speech recognition models combine transcript text with speaker labels and timestamps, so recordings retain information about who spoke and when.
Voxtral is Mistral AI's open-source audio and text model for developers building self-hosted speech applications. It can answer questions about recordings and produce structured summaries within the same model that transcribes speech. Mistral's separate mistral-inference library is archived and no longer maintained; Voxtral supports vLLM and Hugging Face Transformers.
Distil-Whisper is a family of local speech recognition models for developers building English transcription into their apps or services. It reduces Whisper's size and processing time while retaining much of its transcription accuracy. It supports English only.
Whisper is an open source speech recognition model for people who want to transcribe audio on their own hardware. It suits developers adding voice features to an app and anyone working with recordings in multiple languages. OpenAI publishes the models and inference code under the MIT license, so audio can stay on the machine running them.
Moonshine is an on-device AI toolkit for developers building voice agents and applications that listen and speak. It combines speech to text, intent recognition and text to speech in one library. Voice processing stays on the device, and you don't need an account or API keys.