FunASR is a local speech-to-text toolkit for developers building transcription services, voice apps, or audio-processing pipelines. It handles recorded files and live speech, with separate model choices for multilingual transcription, speaker labels, and emotion detection. You can run inference on your own hardware or a self-hosted server.
Fun-ASR-Nano covers Chinese, English, Japanese, and Chinese dialects and accents. Fun-ASR-MLT-Nano is a separate choice for broader language coverage. SenseVoiceSmall transcribes Chinese, English, Japanese, Korean, and Cantonese and adds emotion and audio-event tags. Other supported models include Qwen3-ASR and Whisper. Paraformer-zh-streaming handles live transcription.
Pipelines can combine speech detection, punctuation, timestamps, and speaker separation. SenseVoiceSmall with FSMN-VAD and CAM++ assigns anonymous speaker labels within a recording. The OpenMOSS MOSS-Transcribe-Diarize adapter produces text, timestamps, and speaker labels together for recorded audio. These labels don't identify known people. Batch transcription and SRT subtitle output are also available.
CPU and NVIDIA GPU execution are supported, with vLLM for Nano batch processing. A llama.cpp runtime uses GGUF models on Linux, macOS, Windows, and edge devices without Python at runtime. Docker streaming services, an OpenAI-compatible API, and MCP serving connect transcription to other applications, including Claude and Cursor. Capabilities vary by model and runtime. The open-source toolkit uses the MIT license; model weights have separate licenses.
Claim this page with an email at funasr.com. FunASR gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find FunASR?Promote it
Something wrong or outdated on this page?
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
5.8KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend
10.6KUpdated 21 hours agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
7.7KUpdated 22 hours agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
15.1KUpdated 1 week agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Multilingual#ONNX#Speaker diarization
sherpa-onnx is an open-source toolkit for developers building speech features that run on-device without an internet connection. It handles live microphone transcription and recorded audio, as well as text-to-speech and speaker analysis. Audio processing stays local.
Lemonade is an open source local AI server for people who want to use models on their own hardware or connect them to apps and agents. It handles chat, coding, image generation, speech, transcription, and embeddings. A built-in interface lets you use those capabilities directly, while its server makes them available to other software.
llama-cpp-python brings llama.cpp model inference into Python applications and exposes it through a self-hosted OpenAI-compatible server. It's for developers building local AI applications or connecting existing API clients to models on their own hardware. The package is open source under the MIT license.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.