
MMAudio generates audio that matches a video's action and timing, with text prompts available to guide the result. It also creates audio from text alone. It's for video creators who need sound for silent footage and researchers working on audio generation.
You can run the Python and PyTorch software locally on your own GPU. The project has been tested on Ubuntu, and reported inference uses about 6GB of GPU memory in 16-bit mode. A Gradio interface supports both video and text inputs. Hosted demos are available through Hugging Face, Colab and Replicate; those run outside your local setup.
Its distinguishing feature is a synchronization module that aligns generated sound with video frames. The model trains on both paired audio and video data and paired audio and text data, so it can draw on both kinds of material. Examples include audio generated for footage from Sora, Veo 2 and Hunyuan Video, though these are examples of input footage rather than direct integrations.
The model trains on eight-second clips. It can generate shorter or longer audio, but quality may drop when the duration differs substantially. Human speech can sound unintelligible, incidental background music may be poor, and unfamiliar sound concepts can be difficult for it to reproduce.
The code is open source under the MIT license. Training and evaluation code are included. The pretrained models use datasets with separate licenses, and the project doesn't guarantee their suitability for commercial use.
Claim this page with an email at hkchengrex.com. MMAudio gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find MMAudio?Promote it
Something wrong or outdated on this page?
2.8KUpdated 9 months agoApache-2.0
Windows · Linux#Batch processing#ONNX#Voice activity detection
openWakeWord is a Python library for developers building voice interfaces that listen locally for a chosen word or phrase. It includes English models for triggers such as "hey jarvis" and "alexa", plus phrases for weather and timers. The code uses Apache 2.0. Included pretrained models use CC-BY-NC-SA-4.0, which restricts commercial use.
3.9KUpdated 4 months agoMIT
Web#Hugging Face integration
Stable Audio Tools is an MIT-licensed Python toolkit for developers and audio researchers who want to generate audio on their own hardware or train models on their own recordings. It combines model inference with training and fine-tuning, so you can work with pretrained models or build a model around a specific audio dataset.
10.6KUpdated 3 months agoMIT
#Hugging Face integration#Speaker diarization#Voice activity detection
pyannote.audio is a Python toolkit that separates an audio recording into timed segments labeled by speaker. It's for developers and researchers who need to track who spoke when, with pretrained models that run on their own hardware. The toolkit is open source under the MIT license.