Favicon of Bark

Bark

An open-source text-to-audio model that runs locally on CPU or NVIDIA GPU, with multilingual speech, voice presets and an MIT license.

Bark is Suno's local text-to-audio model for developers and researchers who want to generate speech alongside other sounds. It can produce laughter, crying, music and background noise within its output. Its generative approach suits audio experiments, though it can depart from the supplied script and doesn't guarantee clean, studio-quality speech.

It detects the language from your text and supports English, German, Spanish, French, Hindi, Italian, Japanese, Korean, Polish, Portuguese, Russian, Turkish and simplified Chinese. English is its strongest language. Text that mixes languages can prompt changes in accent. Speaker presets guide tone, pitch, emotion and speech rhythm, and Bark can also generate random voices. It doesn't support custom voice cloning.

Bark is free and open source under the MIT license, with pretrained models available for commercial use. Audio generation runs on your own hardware through Python or Hugging Face Transformers. Hugging Face supplies the model downloads. It works on CPU and CUDA-compatible NVIDIA GPUs, though CPU generation can be considerably slower.

The full model needs about 12 GB of VRAM when kept entirely on the GPU; smaller models fit in about 8 GB and trade some quality for speed. CPU offloading allows smaller GPUs to work too. Bark uses a GPT-style architecture with EnCodec audio representations to generate sound directly from text. Individual generations are typically around 13 seconds long, and the project includes notebook examples for longer audio and more consistent voices.

Similar to Bark