Bark is Suno's local text-to-audio model for developers and researchers who want to generate speech alongside other sounds. It can produce laughter, crying, music and background noise within its output. Its generative approach suits audio experiments, though it can depart from the supplied script and doesn't guarantee clean, studio-quality speech.
It detects the language from your text and supports English, German, Spanish, French, Hindi, Italian, Japanese, Korean, Polish, Portuguese, Russian, Turkish and simplified Chinese. English is its strongest language. Text that mixes languages can prompt changes in accent. Speaker presets guide tone, pitch, emotion and speech rhythm, and Bark can also generate random voices. It doesn't support custom voice cloning.
Bark is free and open source under the MIT license, with pretrained models available for commercial use. Audio generation runs on your own hardware through Python or Hugging Face Transformers. Hugging Face supplies the model downloads. It works on CPU and CUDA-compatible NVIDIA GPUs, though CPU generation can be considerably slower.
The full model needs about 12 GB of VRAM when kept entirely on the GPU; smaller models fit in about 8 GB and trade some quality for speed. CPU offloading allows smaller GPUs to work too. Bark uses a GPT-style architecture with EnCodec audio representations to generate sound directly from text. Individual generations are typically around 13 seconds long, and the project includes notebook examples for longer audio and more consistent voices.
Claim this page and we'll verify you by hand. Bark gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Bark?Promote it
Something wrong or outdated on this page?
2.6KUpdated 2 years ago
macOS · Linux · Web#Batch processing#Hugging Face integration
AudioLDM 2 generates sound effects, music and speech on your own hardware. It's a Python tool for people experimenting with synthetic audio, including sound designers and researchers who want to work with pretrained models. A Gradio browser interface and command-line tools provide access to local generation; a hosted Hugging Face demo is also available.
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
8.4KUpdated 4 months agoApache-2.0
#Batch processing#Hugging Face integration#Multilingual
24.2KUpdated 1 day ago
Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input
6.3KUpdated 10 months agoApache-2.0
#Hugging Face integration#llama.cpp backend#LoRA
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
Higgs Audio is a family of text-to-speech models from Boson AI for developers building narration and conversational audio. Higgs TTS 2 can adapt pacing and intonation to the text and generate dialogue with distinct speakers across multiple languages.
IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.
Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.