Favicon of FunASR

FunASR

Open-source speech-to-text toolkit for offline and self-hosted transcription, with CPU and GPU runtimes, speaker labels, and an OpenAI-compatible API.

FunASR is a local speech-to-text toolkit for developers building transcription services, voice apps, or audio-processing pipelines. It handles recorded files and live speech, with separate model choices for multilingual transcription, speaker labels, and emotion detection. You can run inference on your own hardware or a self-hosted server.

Fun-ASR-Nano covers Chinese, English, Japanese, and Chinese dialects and accents. Fun-ASR-MLT-Nano is a separate choice for broader language coverage. SenseVoiceSmall transcribes Chinese, English, Japanese, Korean, and Cantonese and adds emotion and audio-event tags. Other supported models include Qwen3-ASR and Whisper. Paraformer-zh-streaming handles live transcription.

Pipelines can combine speech detection, punctuation, timestamps, and speaker separation. SenseVoiceSmall with FSMN-VAD and CAM++ assigns anonymous speaker labels within a recording. The OpenMOSS MOSS-Transcribe-Diarize adapter produces text, timestamps, and speaker labels together for recorded audio. These labels don't identify known people. Batch transcription and SRT subtitle output are also available.

CPU and NVIDIA GPU execution are supported, with vLLM for Nano batch processing. A llama.cpp runtime uses GGUF models on Linux, macOS, Windows, and edge devices without Python at runtime. Docker streaming services, an OpenAI-compatible API, and MCP serving connect transcription to other applications, including Claude and Cursor. Capabilities vary by model and runtime. The open-source toolkit uses the MIT license; model weights have separate licenses.

Similar to FunASR