Favicon of MusicGen (AudioCraft)

MusicGen (AudioCraft)

A text-to-music model with melody conditioning and local GPU inference. AudioCraft code is MIT licensed; pretrained weights have a noncommercial license.

Screenshot of MusicGen (AudioCraft) website

MusicGen is Meta AI's music generation model within AudioCraft, a PyTorch library for developers and audio researchers who want to generate music in their own computing environment. It creates music from text descriptions and can use a melody to guide the result. AudioCraft includes both inference and training code, so it's suited to people building audio tools or studying music generation.

Text prompts control the music. Melody conditioning adds a musical reference alongside the written description, giving users another way to shape the output. The library also includes MusicGen Style for generation guided by text and style, plus JASCO for music conditioned on chords, melodies and drum tracks. These are separate models within the same audio research toolkit.

MusicGen uses EnCodec to represent audio as compressed tokens, which a single autoregressive language model generates before the codec decodes them into a waveform. For researchers, the appeal is access to that model architecture together with the code needed to train and run it. AudioCraft provides reusable PyTorch components and training pipelines for developing audio generation models.

The surrounding library covers sound generation as well as music: AudioGen produces environmental sounds from text, and MAGNeT handles text-to-music and text-to-sound generation. It also includes an EnCodec-compatible diffusion decoder and AudioSeal for audio watermarking.

AudioCraft code uses the MIT license. The pretrained MusicGen weights use CC BY-NC 4.0, which restricts their use to noncommercial purposes. Local inference requires a GPU; memory needs vary by model size and output length.

Similar to MusicGen (AudioCraft)