Favicon of Voxtral

Voxtral

An open-source audio AI model for local transcription, translation and Q&A. Runs offline with vLLM or Transformers under Apache 2.0.

Screenshot of Voxtral website

Voxtral is Mistral AI's open-source audio and text model for developers building self-hosted speech applications. It can answer questions about recordings and produce structured summaries within the same model that transcribes speech. Mistral's separate mistral-inference library is archived and no longer maintained; Voxtral supports vLLM and Hugging Face Transformers.

The model combines audio understanding with the text capabilities of Mistral Small. A dedicated transcription mode detects the spoken language automatically, while audio conversations can include several recordings and follow-up questions. It also handles translation. Supported languages include English, Spanish, French, Portuguese, Hindi, German, Dutch and Italian.

For applications that need to work with longer recordings, Voxtral handles up to 30 minutes of audio for transcription or 40 minutes for understanding. You can ask it to compare recordings, summarize their content or respond to a spoken question without connecting a separate speech recognition model to a language model. It also accepts text-only conversations.

Voxtral Small runs offline on your own hardware or behind a server you control, so local audio processing doesn't require sending recordings to a cloud service. Its weights use the Safetensors format and carry the Apache 2.0 license. The Small model needs about 55 GB of GPU memory at bf16 or fp16 precision.

Similar to Voxtral