
AudioLDM 2 generates sound effects, music and speech on your own hardware. It's a Python tool for people experimenting with synthetic audio, including sound designers and researchers who want to work with pretrained models. A Gradio browser interface and command-line tools provide access to local generation; a hosted Hugging Face demo is also available.
Text prompts can describe a sound scene or musical style. For speech, you supply the words and a description of the speaker. The available checkpoints include general sound and music models, a dedicated music model, and speech models trained on GigaSpeech or LJSpeech. The repository also supports audio super-resolution and inpainting.
Hardware support covers CPU, NVIDIA GPUs through CUDA, and Apple Silicon through MPS. The MPS implementation needs about 20 GB of RAM. Output quality can vary with the hardware and random seed, so results may need several attempts.
Its shared audio representation, built with AudioMAE, lets the same learning framework cover speech, music and sound effects. GPT-2 predicts that representation, and a latent diffusion model generates the audio. The research demonstrations also include image-to-audio generation and speech continuation using a short audio sample as context.
Hugging Face Diffusers provides another way to run the models locally. That implementation supports audio of arbitrary length and runs faster than the native implementation, with official checkpoints available through the Hugging Face Hub. The repository is licensed under CC BY-NC-SA 4.0, and the published AudioLDM2 weights also carry noncommercial terms; do not treat the availability of local inference as permission for commercial use.
Claim this page and we'll verify you by hand. AudioLDM 2 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find AudioLDM 2?Promote it
Something wrong or outdated on this page?
39.3KUpdated 2 years agoMIT
#Hugging Face integration#Multilingual
Bark is Suno's local text-to-audio model for developers and researchers who want to generate speech alongside other sounds. It can produce laughter, crying, music and background noise within its output. Its generative approach suits audio experiments, though it can depart from the supplied script and doesn't guarantee clean, studio-quality speech.
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
9.7KUpdated 1 day ago
macOS · Windows · Linux · Docker · Web#Batch processing#ControlNet#GGUF
7.2KUpdated 2 years agoApache-2.0
macOS · Linux · Docker · Web#Multilingual#Voice cloning
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
2.3KUpdated 10 months agoApache-2.0
macOS · Windows · Linux · Docker#Hugging Face integration#Multimodal input
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.
Wan2GP brings video, image, music and speech generation to your own computer, with particular attention to GPUs with limited memory. It's for creators who want several media models in one browser interface. The project builds on Wan-Video/Wan2.1.
Zonos is an open-source text-to-speech model for people who want to generate speech and clone voices on their own hardware. It can match a speaker from a short reference recording, with controls for delivery and emotion. The code uses the Apache 2.0 license.
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
DiffRhythm is a local AI music generation model for musicians, developers and researchers who want to create full-length songs on their own hardware. It uses latent diffusion to generate songs with vocals and accompaniment, and can also produce instrumental music. The full model supports songs up to 4 minutes and 45 seconds.