
MusicGen is Meta AI's music generation model within AudioCraft, a PyTorch library for developers and audio researchers who want to generate music in their own computing environment. It creates music from text descriptions and can use a melody to guide the result. AudioCraft includes both inference and training code, so it's suited to people building audio tools or studying music generation.
Text prompts control the music. Melody conditioning adds a musical reference alongside the written description, giving users another way to shape the output. The library also includes MusicGen Style for generation guided by text and style, plus JASCO for music conditioned on chords, melodies and drum tracks. These are separate models within the same audio research toolkit.
MusicGen uses EnCodec to represent audio as compressed tokens, which a single autoregressive language model generates before the codec decodes them into a waveform. For researchers, the appeal is access to that model architecture together with the code needed to train and run it. AudioCraft provides reusable PyTorch components and training pipelines for developing audio generation models.
The surrounding library covers sound generation as well as music: AudioGen produces environmental sounds from text, and MAGNeT handles text-to-music and text-to-sound generation. It also includes an EnCodec-compatible diffusion decoder and AudioSeal for audio watermarking.
AudioCraft code uses the MIT license. The pretrained MusicGen weights use CC BY-NC 4.0, which restricts their use to noncommercial purposes. Local inference requires a GPU; memory needs vary by model size and output length.
Claim this page with an email at audiocraft.metademolab.com. MusicGen (AudioCraft) gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find MusicGen (AudioCraft)?Promote it
Something wrong or outdated on this page?
1.8KUpdated 2 months agoApache-2.0
macOS#MLX#Streaming inference
Magenta RealTime 2 is a local AI music model and synthesis engine for musicians and developers who want to play or build AI musical instruments on a laptop. It generates streaming audio in real time, with open weights and code under the Apache 2.0 license.
21.1KUpdated 2 days agoApache-2.0
macOS · Web#GGUF#Hugging Face integration#Multilingual
2.3KUpdated 10 months agoApache-2.0
macOS · Windows · Linux · Docker#Hugging Face integration#Multimodal input
13.2KUpdated 2 days agoApache-2.0
#ControlNet#Image-to-image#Inpainting
9.2KUpdated 6 days agoApache-2.0
Docker#Image-to-image#Inpainting#Multimodal input
14.2KUpdated 5 days ago
#Hugging Face integration#Multimodal input
OpenCLIP is a Python and PyTorch library for developers and researchers who want to match images with text on their own hardware. It implements OpenAI's CLIP approach: images and descriptions become numerical representations that the model can compare. This supports image search and zero-shot classification, where text labels define the categories without a separate classifier trained for each task.
Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.
DiffRhythm is a local AI music generation model for musicians, developers and researchers who want to create full-length songs on their own hardware. It uses latent diffusion to generate songs with vocals and accompaniment, and can also produce instrumental music. The full model supports songs up to 4 minutes and 45 seconds.
DiffSynth-Studio is a Python diffusion model engine for developers and researchers who want to generate media and train models on their own hardware. It supports large models on consumer GPUs through memory offloading and quantization, with inference and training in the same framework. It's open source under Apache 2.0.
ModelScope combines a hosted model and dataset hub with a Python library you can run locally. It's for developers and researchers who want to use AI models in their own applications, fine-tune them on their own data, or compare their performance. The library is open source under Apache 2.0.