Moonshine is an on-device AI toolkit for developers building voice agents and applications that listen and speak. It combines speech to text, intent recognition and text to speech in one library. Voice processing stays on the device, and you don't need an account or API keys.
Its streaming approach processes speech while someone is still talking, reducing the delay before an application can respond. That makes it relevant to conversational interfaces where waiting for a complete recording would interrupt the exchange. Developers can also use it for transcription without building a full voice agent.
The speech recognition models are trained from scratch. The range includes models focused on accuracy and small models suited to constrained devices, so developers can choose according to the application and hardware. The documentation also covers model customization for particular domains.
Moonshine supports macOS, Windows and Linux, as well as iOS, Android and Raspberry Pi. Python and JavaScript/WASM support let developers use the toolkit across different application environments, including browsers. A C API is available for integration into native software.
The toolkit is open source under the MIT License. Its models use MIT by default across languages and sizes, with a specific exception: legacy non-streaming models for languages other than English use the non-commercial Moonshine Community License.
Claim this page and we'll verify you by hand. Moonshine gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Moonshine?Promote it
Something wrong or outdated on this page?
54.5KUpdated 4 weeks agoMIT
#Hugging Face integration#Multilingual#Quantization
VibeVoice is a family of MIT-licensed, open-source voice AI models for developers and researchers building local transcription or speech generation tools. Its speech recognition models combine transcript text with speaker labels and timestamps, so recordings retain information about who spoke and when.
382Updated 3 years agoMIT
#Home Assistant integration#Multilingual#Voice activity detection
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
31.3KUpdated 3 weeks agoMIT
macOS · Windows · Linux#Ollama integration#OpenAI-compatible API#Streaming inference
11.2KUpdated 5 months agoApache-2.0
macOS · iOS · Web#Hugging Face integration#MLX#Quantization
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
Rhasspy 3 is an early developer-preview local voice assistant toolkit for developers building their own assistants or adding voice control to Home Assistant. The project is archived and no longer maintained. It keeps data on your computer unless you choose to send it elsewhere, and its speech components support languages beyond English.
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
Meetily is a local AI meeting assistant for people who want meeting notes while keeping recordings on their own device. It captures calls from Zoom, Google Meet, Microsoft Teams and other meeting software without placing a bot in the meeting. You can watch the transcript appear during the call, then generate a summary.
Moshi is a voice AI model and dialogue framework that can listen while it speaks. It processes speech directly, retaining information such as emotion and non-verbal cues that a text transcription can miss. It's aimed at researchers and developers building spoken AI applications, with local inference and self-hosted server options.
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.