Favicon of VibeVoice

VibeVoice

Open-source voice AI models for local transcription and speech generation, with MIT licensing, CPU inference, and streaming audio support.

VibeVoice is a family of MIT-licensed, open-source voice AI models for developers and researchers building local transcription or speech generation tools. Its speech recognition models combine transcript text with speaker labels and timestamps, so recordings retain information about who spoke and when.

VibeVoice-ASR can process an hour of audio in one pass. That whole-recording context helps it track speakers across long conversations. It supports multilingual transcription and custom hotwords for names, specialist vocabulary, or background context. A separate streaming model produces transcripts as speech arrives, with speaker tracking and hotword support.

For CPU deployment, VibeVoice-ASR-BitNet provides real-time recognition without a GPU. Developers can also use Hugging Face Transformers or serve recognition through vLLM, and the project includes code for fine-tuning ASR. Microsoft Foundry and online playgrounds offer hosted ways to try recognition; the CPU inference engine runs on your own hardware.

VibeVoice-Realtime accepts text as it arrives and starts generating speech before the full input is available. It supports extended speech output, making it relevant to applications that need spoken responses to gradually generated text. The speech generation architecture combines a language model for textual context with diffusion-based audio generation.

The project targets research and development. The original long-form, multi-speaker VibeVoice-TTS code has been removed from the repository, while VibeVoice-Realtime remains available.

Similar to VibeVoice