VibeVoice is a family of MIT-licensed, open-source voice AI models for developers and researchers building local transcription or speech generation tools. Its speech recognition models combine transcript text with speaker labels and timestamps, so recordings retain information about who spoke and when.
VibeVoice-ASR can process an hour of audio in one pass. That whole-recording context helps it track speakers across long conversations. It supports multilingual transcription and custom hotwords for names, specialist vocabulary, or background context. A separate streaming model produces transcripts as speech arrives, with speaker tracking and hotword support.
For CPU deployment, VibeVoice-ASR-BitNet provides real-time recognition without a GPU. Developers can also use Hugging Face Transformers or serve recognition through vLLM, and the project includes code for fine-tuning ASR. Microsoft Foundry and online playgrounds offer hosted ways to try recognition; the CPU inference engine runs on your own hardware.
VibeVoice-Realtime accepts text as it arrives and starts generating speech before the full input is available. It supports extended speech output, making it relevant to applications that need spoken responses to gradually generated text. The speech generation architecture combines a language model for textual context with diffusion-based audio generation.
The project targets research and development. The original long-form, multi-speaker VibeVoice-TTS code has been removed from the repository, while VibeVoice-Realtime remains available.
Claim this page and we'll verify you by hand. VibeVoice gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find VibeVoice?Promote it
Something wrong or outdated on this page?
11.2KUpdated 1 month ago
macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference
Moonshine is an on-device AI toolkit for developers building voice agents and applications that listen and speak. It combines speech to text, intent recognition and text to speech in one library. Voice processing stays on the device, and you don't need an account or API keys.
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
6.3KUpdated 10 months agoApache-2.0
#Hugging Face integration#llama.cpp backend#LoRA
9.4KUpdated 3 weeks agoMIT
Docker#Batch processing#GGUF#Hugging Face integration
39.3KUpdated 2 years agoMIT
#Hugging Face integration#Multilingual
Bark is Suno's local text-to-audio model for developers and researchers who want to generate speech alongside other sounds. It can produce laughter, crying, music and background noise within its output. Its generative approach suits audio experiments, though it can depart from the supplied script and doesn't guarantee clean, studio-quality speech.
32.9KUpdated 2 weeks ago
#Batch processing#Multilingual#Multimodal input
Fish Speech, currently featuring Fish Audio S2 Pro, is a self-hosted text-to-speech system for creators producing narration and developers building voice applications. It combines voice cloning with control over emotion and delivery within a script. Code and model weights use the custom FISH AUDIO RESEARCH LICENSE.
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.
SenseVoice is a local speech recognition model that adds language, emotion and sound-event tags to transcriptions. It's for developers building voice applications or analyzing recordings on their own hardware, particularly those working with Mandarin and Cantonese. The project is open source under the MIT license.