Favicon of AudioLDM 2

AudioLDM 2

A local AI audio generator that turns text into sound effects, music and speech. Runs on CPU, NVIDIA CUDA or Apple Silicon with Hugging Face Diffusers support.

Screenshot of AudioLDM 2 website

AudioLDM 2 generates sound effects, music and speech on your own hardware. It's a Python tool for people experimenting with synthetic audio, including sound designers and researchers who want to work with pretrained models. A Gradio browser interface and command-line tools provide access to local generation; a hosted Hugging Face demo is also available.

Text prompts can describe a sound scene or musical style. For speech, you supply the words and a description of the speaker. The available checkpoints include general sound and music models, a dedicated music model, and speech models trained on GigaSpeech or LJSpeech. The repository also supports audio super-resolution and inpainting.

Hardware support covers CPU, NVIDIA GPUs through CUDA, and Apple Silicon through MPS. The MPS implementation needs about 20 GB of RAM. Output quality can vary with the hardware and random seed, so results may need several attempts.

Its shared audio representation, built with AudioMAE, lets the same learning framework cover speech, music and sound effects. GPT-2 predicts that representation, and a latent diffusion model generates the audio. The research demonstrations also include image-to-audio generation and speech continuation using a short audio sample as context.

Hugging Face Diffusers provides another way to run the models locally. That implementation supports audio of arbitrary length and runs faster than the native implementation, with official checkpoints available through the Hugging Face Hub. The repository is licensed under CC BY-NC-SA 4.0, and the published AudioLDM2 weights also carry noncommercial terms; do not treat the availability of local inference as permission for commercial use.

Similar to AudioLDM 2