Favicon of MMAudio

MMAudio

Open-source video-to-audio and text-to-audio software you can run locally on a GPU. Uses an MIT license; tested on Ubuntu.

Screenshot of MMAudio website

MMAudio generates audio that matches a video's action and timing, with text prompts available to guide the result. It also creates audio from text alone. It's for video creators who need sound for silent footage and researchers working on audio generation.

You can run the Python and PyTorch software locally on your own GPU. The project has been tested on Ubuntu, and reported inference uses about 6GB of GPU memory in 16-bit mode. A Gradio interface supports both video and text inputs. Hosted demos are available through Hugging Face, Colab and Replicate; those run outside your local setup.

Its distinguishing feature is a synchronization module that aligns generated sound with video frames. The model trains on both paired audio and video data and paired audio and text data, so it can draw on both kinds of material. Examples include audio generated for footage from Sora, Veo 2 and Hunyuan Video, though these are examples of input footage rather than direct integrations.

The model trains on eight-second clips. It can generate shorter or longer audio, but quality may drop when the duration differs substantially. Human speech can sound unintelligible, incidental background music may be poor, and unfamiliar sound concepts can be difficult for it to reproduce.

The code is open source under the MIT license. Training and evaluation code are included. The pretrained models use datasets with separate licenses, and the project doesn't guarantee their suitability for commercial use.

Similar to MMAudio